Pith. sign in

Paper Citation Record · LEDGER

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2508.20096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20096 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:19:50.983692Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T02:18:01.817666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:23:30.927676Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f766a68c-d974-4add-b0aa-9fda25056b8a · outbound

This paper cites Agent S: An Open Agentic Framework that Uses Computers Like a Human.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agent S: An Open Agentic Framework that Uses Computers Like a Human

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.504801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.504801Z digest=sha256:b5f6b6f32469e1256a399b252d4b129dda2fb7fd354e7954917a8d62d506bb25

Observation 536767b3-0f6a-43fb-9d37-1b179687ffb0 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.510962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.510962Z digest=sha256:f7d0276df59a5ab8d1f3aaa67791ebf576acb34ce77bc17658a1fba3c05dd239

Observation 406a4475-a1ac-43d3-972a-32b254eb6b28 · outbound

This paper cites Claude computer use.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Claude computer use

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.250165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.516921Z digest=sha256:5c4e860a0584e73f5bc2f4b9fdd660768c7ad63ff86417ee1b12dc749ddf95d2

Observation f339c29b-e9a9-4cff-8cf7-2ef964d18774 · outbound

This paper cites Claude’s extended thinking.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Claude’s extended thinking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.228819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.523515Z digest=sha256:f0864ab1d6648acd87e9cc77beb5356b56914986de3c8ebbb117fadb76240fb5

Observation b6a231e3-0f55-443c-aef9-2782a0346cb4 · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.205749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.529290Z digest=sha256:563ec16fb4c21680bc2d9e99ea59bb8eeff6339063e914b9706b1d9bd01fee9a

Observation 6cb99d2e-0156-49e3-bd19-399f330b3ddb · outbound

This paper cites Qwen2.5-VL Technical Report.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.536624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.536624Z digest=sha256:0ca6113cf2d2170a0756dc6123af9830a9f07c3e54749c99afa539a53764acbd

Observation a38b0a0b-03d4-4aaf-a8b6-f22418e0e782 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Grounding large language models in interactive environments with online reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.178064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.543496Z digest=sha256:0966f506e5aa57e678a1c9031da071323d9f0d0e3a07fcdfe93bf850939d350f

Observation 64a1b265-369f-4a1a-a4fb-9f123e405ae7 · outbound

This paper cites Bail: Best-action imitation learning for batch deep reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Bail: Best-action imitation learning for batch deep reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.152163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.549403Z digest=sha256:1af2a14f325137d91611748f09fe60301da4ce9e3729ce2d8599fe24c19de24d

Observation 2622bfbe-7249-4901-b021-73b11b1d5df7 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.556970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.556970Z digest=sha256:323392d3c37f85911d7d37ff9ad0d4e87730ba58f92a25a107f291ae902105a0

Observation 65179828-9feb-452d-bf4d-2197d681c635 · outbound

This paper cites Neuroplasticity.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Neuroplasticity

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.123125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.566644Z digest=sha256:14857f8ba791650aea632e301e59005e62b0d9dd9a6800448cec4118cfa5b71a

Observation e5a6931a-e40b-49fe-afd6-f6fa08abcc63 · outbound

This paper cites MM-IFEngine: Towards Multimodal Instruction Following.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning MM-IFEngine: Towards Multimodal Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.573543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.573543Z digest=sha256:6a9ac3585c7e1b6581be39a9a43863a2ec90116bf755ddba85b060690116fa71

Observation 8a2d30db-18ea-47e0-a0b7-27dae693c691 · outbound

This paper cites Gemini 2.5 Pro Preview (03-25).

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Gemini 2.5 Pro Preview (03-25)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.095948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.582712Z digest=sha256:adb266175bbb7c1945f6efe249e4f6aff044653e1f4c77452a3ab0a7e842d744

Observation 31deaa7f-81d4-4e88-8040-767a42a2ee90 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.588175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.588175Z digest=sha256:acbb1abf2705107d4d1c4faf07417e84dc9148cf8eeac1b8d1cc645a1a7cb4e3

Observation ed1dd05c-ad83-499f-881b-79ee669eb433 · outbound

This paper cites The Llama 3 Herd of Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.596218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.596218Z digest=sha256:517b592afc9528141f7a127c60dde0d30aac1ed2c042240bec7a85e370cf7e10

Observation 45d9ce59-dda3-43cc-b9af-960c5bcf32b1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.606852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.606852Z digest=sha256:70860bace651b810d4d19f3a61cebe940e1120b1ba8a95ebd6851e4a76c8c83d

Observation edf07ee9-c9df-4a15-a4be-6b583c55644a · outbound

This paper cites Neuroplasticity and rehabilitation.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Neuroplasticity and rehabilitation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.070739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.615498Z digest=sha256:61bb249960422d8c6ecf3d644518d0294b3d9fcf093048e7ebb7f4bcdeaf36b1

Observation 02aadcec-4134-4e6f-8cd2-42125ae232c6 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.623148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.623148Z digest=sha256:07a6a8e66645e99f71e3c9ea22b5bd2351c456102cf5f2c6abfb662046206650

Observation 0f880a2d-6315-4c62-ae82-8b20b54c5721 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.628996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.628996Z digest=sha256:8d7b2a9a419fc895ae89e02fe302041a3947b517bb37b178bd38ec82ad783dca

Observation a6169206-4609-4bd3-bbd1-aec2b95a5de2 · outbound

This paper cites Cogagent: A visual language model for gui agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Cogagent: A visual language model for gui agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.635778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.635778Z digest=sha256:d04b0be0183346d0cfba16b51e7ff8e213cd1b785c81bd2b3637c2843a41f29e

Observation 487972e9-9ee5-409a-80f6-57864f84b357 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Lora: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.642941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.642941Z digest=sha256:e41b8055d905e8aba956709c53f480fa308fcdaa2cd6a69aa38e57f2c9436e95

Observation 831f7ad4-deb4-4d3f-83a0-04b04207cbdf · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.648306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.648306Z digest=sha256:03a8e305a15cf9c733a18bfc91e69315e592cca3ef276b5dd6af9798eebd3623

Observation ed5549b3-3d54-4fb2-9815-d06a63b473b4 · outbound

This paper cites Os agents: A survey on mllm-based agents for general computing devices use, 2024 b.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Os agents: A survey on mllm-based agents for general computing devices use, 2024 b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.014453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.653801Z digest=sha256:bb5cd0e36b74f7cc1918af6c716540adc63574302f2403c8eb15c2dd1bc7cfa2

Observation 5d76df08-9209-4cf1-9509-330a6d06bb1f · outbound

This paper cites Mechanisms of motor learning in the cerebellum.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Mechanisms of motor learning in the cerebellum

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.990888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.659738Z digest=sha256:a363a05f3b2643fed13b0d41619ba8ea6dd587f79a90dd59a62e50b83c30754d

Observation e50f27c3-bb7f-4721-a431-21dfb23fabf9 · outbound

This paper cites Autowebglm: A large language model-based web navigating agent.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Autowebglm: A large language model-based web navigating agent

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.972155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.666070Z digest=sha256:7623ebc0520393923f4bbd3ece94a1f4b2d8e5b43929180bfa4cfda176fec770

Observation 96710486-0104-4b5f-8a91-a545d55cb11b · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.672939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.672939Z digest=sha256:7ac0c21dfaf7bcf9cc4bc07dee9d997dae9f9f78b20fe8b071f867bce2cd860b

Observation 312fc508-ed40-4cf0-b844-da993d00925c · outbound

This paper cites Visual instruction tuning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.950298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.679108Z digest=sha256:74b6df5b8f998ac64fbc46aba53e9a25648280192646fc52d4eee5dc3729eb73

Observation 18e22ced-d723-44c6-9338-3c010a917d43 · outbound

This paper cites BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.685669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.685669Z digest=sha256:f356afebb238447addd0272e3496b0116f382f6489d97dbc80b9a186d9e3208e

Observation ecbd4669-5baa-47a4-8372-5898edb01355 · outbound

This paper cites Agentrewardbench: Evaluating automatic evaluations of web agent trajectories.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agentrewardbench: Evaluating automatic evaluations of web agent trajectories

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.694654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.694654Z digest=sha256:9d4194d6ddf7feb8f9a4f96c5e96d8948ec0e8d31d2cec527e2297ca21d0ba35

Observation bb49802b-a5cc-49ca-85ed-62f31cdaddfb · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.700711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.700711Z digest=sha256:1ec62b665be8a8cc4ffc5803a466089ddce894da3f9d5b2b646aeb30245f012a

Observation a30da84f-5130-42c3-b14c-0118dbad65fe · outbound

This paper cites Gui agents: A survey.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Gui agents: A survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.708341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.708341Z digest=sha256:0100760cc3387b890828a8647c705babea605bd409f4c69492a2cd6496ed2fd7

Observation 71692cd5-27e2-494b-9579-7bbd73cebc04 · outbound

This paper cites GPT-4 Technical Report.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.713835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.713835Z digest=sha256:3a6eedb41acc6762c6cc811d416a79b89afc5ca574c548afd4ec1934057881ce

Observation bdb4c472-556e-4dc6-9bc3-86bfd27ddd69 · outbound

This paper cites Operator.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Operator

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.918205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.722716Z digest=sha256:0ca5212a202d5908867e63d3f2afd31553043a9cd57d28f6346573e11c6354cb

Observation 01bd980b-836f-4962-949d-b70a2bec033d · outbound

This paper cites Training language models to follow instructions with human feedback.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.730369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.730369Z digest=sha256:a1bac33fc41966df51f230ffe1c551b53d7ffe6ebd33106d5e1df1c2db5d517d

Observation 24c293cd-f165-4b49-a0e8-e7ca369e0c89 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Autonomous Evaluation and Refinement of Digital Agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.737187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.737187Z digest=sha256:d50c7fa89f8e9f1a3f4b5d10ce88ded7d0e6ebbcaca76f32b380f9b6c11523ab

Observation 057d2df6-59df-499e-a5bc-f172508ff7c4 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.752711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.752711Z digest=sha256:8027f2568d4dda4115de8fd048600d7a6a94410dc92a90efdddfa7b4de26e17c

Observation ceb3ec28-6bb5-463b-97f4-03a7fcc6ba6e · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.759166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.759166Z digest=sha256:17a0274ff30085403303cbd47fa3226276368c245e25843ac7a4daca71cf90e1

Observation a3a602f7-27ae-4246-b411-6d38d2dd4149 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.765656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.765656Z digest=sha256:f90f7bdd1adc26aaac88f14589ee514a9c9760731dffd4a7818120a2e9be92a4

Observation 032941e2-7c06-4749-97ec-d4baeeb78e46 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.774012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.774012Z digest=sha256:691154eff9c60735171a6f130212d1d021ebd5b5ece907fe324cc17f43bc62c8

Observation 2958264b-f2eb-46b2-8a74-dc358820b2c0 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.779222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.779222Z digest=sha256:5373339471234af666f6289689c48ac29aba32e1c18ca91ea09cdd7233d6d855

Observation 68d00cff-0bf1-496a-b6dc-07745af948d0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.785368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.785368Z digest=sha256:ee6672c222fafe214c0894159bdef638a4f00ec192b88ccae1c1c4c527dcac11

Observation 4cbdfc24-f676-4341-af48-f3f580f3ad50 · outbound

This paper cites Coact-1: Computer-using agents with coding as actions.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Coact-1: Computer-using agents with coding as actions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.794782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.794782Z digest=sha256:6439404611bfc4b473d206efb01b05fa3977400064b93d75fe15a285772d70d0

Observation a5e2c418-4155-4d6a-9d03-fa507e6e8482 · outbound

This paper cites A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.801652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.801652Z digest=sha256:c971eefe16d41c9cd719bdb69fe3e54dd79001f44c84e22ed38064ad43f01494

Observation 329bea5a-7f2f-4714-adc0-154d95ba13d9 · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.807916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.807916Z digest=sha256:4e4b4881515f635f30f94943d7c189549c4311ae3c3f0914e0f2c4fd81ba11f2

Observation 40e44535-cd47-4b76-8e5a-ee9bbb0a8970 · outbound

This paper cites ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.814391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.814391Z digest=sha256:587798ba628d0c343438f69d13c4f434d2159711c366ea90042821fb317784ef

Observation 84696db2-b6ed-48cf-9186-844744fabb5a · outbound

This paper cites X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:19:51.797076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.820124Z digest=sha256:c874cfb14b6e6c1ea1b43ff6cec4da6c6f147f1ece43ffe27d83160234d50eed

Observation aedb1683-460c-42d4-b869-397baa43815d · outbound

This paper cites Bootstrap3d: Improving 3d content creation with synthetic data.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Bootstrap3d: Improving 3d content creation with synthetic data

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.861294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.825509Z digest=sha256:e5bca32068208087b467e02abd0be7933786f2f873066628414e8e9f71f9fb30

Observation 49b4daea-61b9-4c3d-bbb7-96c1f24217f2 · outbound

This paper cites SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.834935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.834935Z digest=sha256:dbd02f09eee9565fd53a0d501b44a47f72e72d56660d0506421184f2302495da

Observation da93040c-03ae-4654-8271-0c5ef21148b2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.840176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.840176Z digest=sha256:81803eca57b2c42636729c9fddf2a90e6daf06dedb737ad3566702a503240236

Observation c8f2b661-b8cf-44b7-8ffc-53fecc639b54 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.845468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.845468Z digest=sha256:c7a58ea7a55560d79b8937ba75e80f3f13a1a916f1e52dc93dbcc1bcc482b963

Observation 19680e7c-890b-4461-a2ff-a5fb8342b518 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.851707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.851707Z digest=sha256:5fdc701b8f637d8f442adfaae6e8197b1e2cfc2af004afe0cd80a142b59b1294

Observation df0f4c89-4637-4436-b04f-92e8e56b571b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.858856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.858856Z digest=sha256:d4d31e64ec7570122b7177d445b8b3e0c3f93295d3f9ff31229dc46f7ae5cc96

Observation 797d43fd-7197-41e7-bf5e-968890cc8ce2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.864786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.864786Z digest=sha256:c238795d60d7a878865ff6633fbbb13fcb8f267b1ce9ce05fd8763497815bb2f

Observation 9907cfa2-4638-489e-95d6-861cf9afaee7 · outbound

This paper cites GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.871679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.871679Z digest=sha256:37c9410ea452fc9629fedab7b2fe61e9d3bc90a30f77952f5b7c7ad5f47fe09a

Observation ff81c056-6cc3-41b8-be74-c0357b4352f1 · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.877434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.877434Z digest=sha256:539271fd0ca7a87d0cf8190cf726f0a789c31e9811f9ed9f05722cffccff98ab

Observation cd5029d8-518e-4fcb-b74f-abc604e30118 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.883825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.883825Z digest=sha256:cc90ecf79582c834f127779801cb231d026cb3f1f7b32a82fcfe432ad22bf592

Observation 44df1c76-9e65-4a08-b564-04d25ab53a75 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.803037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.890161Z digest=sha256:ebccb90eefb3487019abbb0df8770f4b304855c9a303e5cbadc671a7a40e5b1a

Observation 1e0ef00a-f26e-4e5a-b5f2-e24d4edeb7f1 · outbound

This paper cites Scaling computer-use grounding via user interface decomposition and synthesis.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Scaling computer-use grounding via user interface decomposition and synthesis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.897174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.897174Z digest=sha256:dfd206e110560153793ed7b1104166d8350f5b0a082b1fadffe136018953d3d1

Observation 909b480f-cc88-4ceb-92a9-a48cf0da6589 · outbound

This paper cites ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.902512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.902512Z digest=sha256:f1bfe372e8ea529587c7e19c583a42433292089b0e91d990c3ca2b2fe0556306

Observation 5157dc22-57cf-42a6-9430-acbbbc010f77 · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.914458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.914458Z digest=sha256:2003df591ea4b8267061ad47f7b166489a7d771db91949bd2782ef6188f94a32

Observation b838e4cb-25af-4756-9ced-7248d9a4dedf · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.775523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.921108Z digest=sha256:a919edb9a745e9494e45955507bab256591707c980e328718e9d74eb2bade352

Observation 2ce47284-8dae-4e6f-a4d4-2670257e4e41 · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Appagent: Multimodal agents as smartphone users

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.753496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.928045Z digest=sha256:77273a3c5440e5047dc10fd089c765297c9f0eee2f2830a878b534d6f9a4a157

Observation a6a42c0f-65c9-4b3d-b61c-f0b09b075a31 · outbound

This paper cites Android in the Zoo: Chain-of-Action-Thought for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Android in the Zoo: Chain-of-Action-Thought for GUI Agents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.933323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.933323Z digest=sha256:e00426e1f7ca65c97cdc4c57f0705ccfd610edc724c0dfeba55c4c2f328e1271

Observation 62386c80-f36b-4f40-9ee2-524a2c2aa569 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.939908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.939908Z digest=sha256:7bf618b808dae15f67e3b8db4019952551c231c45d698742743a8c1db87d2528

Observation 535dcb56-38ea-4662-8ec2-dd8827cf14c3 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.947042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.947042Z digest=sha256:0e099826a09ff5081323786d9ce9846c62af85a91f5c26cd0dd18dc430b27df8

Observation 7aa67b8e-3163-4fd6-bab5-b1873660a265 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.953787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.953787Z digest=sha256:2276c6a1f1209926a736eb2ee979b990ed2bedb9b39c43f56dd642217b425cb6

Observation e39bfaa4-d417-4291-86f9-0218b6453e8f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.959760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.959760Z digest=sha256:7bcfbfd1d3d00eb706cdcdf89e37e8aa7a65953604964b98846dae626d4a2438

Observation 38ef65cc-3db8-4ecf-ac21-78bee8854e5b · outbound

This paper cites write newline.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.966719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.966719Z digest=sha256:77842705c181f7b640e51667eeba6b85321384eea709e13dbe40ff620d398b29

Observation 90d3248b-ada5-4690-9aa5-29236f09ef3a · outbound

This paper cites @esa (Ref.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning @esa (Ref

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.972838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.972838Z digest=sha256:76d104f30c34697c5b1cfca58d30c9a7f5245d8433ef7cf5ebfda2520235c5c1

Observation 74f7dfa6-d1ef-4e47-bc5d-adf7b091d966 · outbound

This paper cites an unresolved cited work.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.978224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.978224Z digest=sha256:1d5273ae94c25c5401dd3a7beb87c5efc76b723d8f6787e8bb9c786607e8ff00

Observation 9ecf09ff-f65d-4172-ad5d-0ed4431b759a · outbound

This paper cites planner" and an.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning planner" and an

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-05T15:19:50.983692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.983692Z digest=sha256:aa1a37c33b0c88bb488a350b8d5f6a223c338ca091b7693ea6221ab8d6603b36

Pith citing papers

Observation 460e88b7-e6db-4b59-946e-d4d804092b43 · inbound

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents cites this paper.

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T14:15:55.180284Z digest=sha256:3c2f18c0f72bb5503abfde29ba6c799415241faabd657d70d5c46c88fc131920

Observation 9d79eac9-d5b5-4c6e-bdd7-ca9363ed2204 · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.817666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.817666Z digest=sha256:f0d444e08dfcb7bffbaa07d75ea4c7af270dc8b52d05f8b847440766c41b8248