Pith. sign in

Paper Citation Record · LEDGER

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 2 inbound Pith citation observations for arXiv:2508.20096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20096 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:19:50.983692Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T02:18:01.817666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T14:23:30.927676Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved51
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f766a68c-d974-4add-b0aa-9fda25056b8a · outbound

This paper cites Agent S: An Open Agentic Framework that Uses Computers Like a Human.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agent S: An Open Agentic Framework that Uses Computers Like a Human

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.504801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.504801Z digest=sha256:f5c226763e84296e0d2ffa76347fb4b91f5c5d88600c939fc7240fd6ac644cb6

Observation 536767b3-0f6a-43fb-9d37-1b179687ffb0 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.510962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.510962Z digest=sha256:6217106cb492fb32aec182ae2ee10141f1a2e6392ba3d7a53abcfc96eb211c78

Observation 406a4475-a1ac-43d3-972a-32b254eb6b28 · outbound

This paper cites Claude computer use.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Claude computer use

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.250165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.516921Z digest=sha256:078faf45e1d23f3bd713b2c6949ab46361017fc58e383b3731c31f3c51454185

Observation f339c29b-e9a9-4cff-8cf7-2ef964d18774 · outbound

This paper cites Claude’s extended thinking.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Claude’s extended thinking

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.228819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.523515Z digest=sha256:5d29a14a01906e32029302ccf9fb5d4dccc908d4dfcc35c1c27a97312c71c305

Observation b6a231e3-0f55-443c-aef9-2782a0346cb4 · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.205749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.529290Z digest=sha256:edcdab13166e93553824d68161740d75219d6c665e6353f710547552ee305423

Observation 6cb99d2e-0156-49e3-bd19-399f330b3ddb · outbound

This paper cites Qwen2.5-VL Technical Report.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.536624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.536624Z digest=sha256:311806fb1c5d12d0abea2122933b820c9b474c8952d7a4e49adbf51623d98e98

Observation a38b0a0b-03d4-4aaf-a8b6-f22418e0e782 · outbound

This paper cites Grounding large language models in interactive environments with online reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Grounding large language models in interactive environments with online reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.178064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.543496Z digest=sha256:8d6a77ccdd02d2f223363e012377cfef2cfb7da4f0ce8d94331eb1863af6dd0b

Observation 64a1b265-369f-4a1a-a4fb-9f123e405ae7 · outbound

This paper cites Bail: Best-action imitation learning for batch deep reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Bail: Best-action imitation learning for batch deep reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.152163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.549403Z digest=sha256:34efe3885dda5c5a0ccc0f7b85d7fc94f6aa07afa19721a71f18ee58e90dbaae

Observation 2622bfbe-7249-4901-b021-73b11b1d5df7 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.556970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.556970Z digest=sha256:b973a56377c121fcd631db372cdaed429ed685ada6b2e72b88fd7032ef28e79e

Observation 65179828-9feb-452d-bf4d-2197d681c635 · outbound

This paper cites Neuroplasticity.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Neuroplasticity

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.123125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.566644Z digest=sha256:7ddbc4bbd6723460a61e5e9c6d20320c5789cc7ebcc4cf024ee40f4f18f4dab0

Observation e5a6931a-e40b-49fe-afd6-f6fa08abcc63 · outbound

This paper cites MM-IFEngine: Towards Multimodal Instruction Following.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning MM-IFEngine: Towards Multimodal Instruction Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.573543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.573543Z digest=sha256:6270a47cca905b75656ecec107e4c23161ba401774d9af65089b47d8a6b89eb4

Observation 8a2d30db-18ea-47e0-a0b7-27dae693c691 · outbound

This paper cites Gemini 2.5 Pro Preview (03-25).

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Gemini 2.5 Pro Preview (03-25)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.095948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.582712Z digest=sha256:dce517bfd0c9fc94dd54eef28df4f23deb9b03f103c7518bca6f7afc3e3d606b

Observation 31deaa7f-81d4-4e88-8040-767a42a2ee90 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.588175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.588175Z digest=sha256:d33a20b2662d96177e423cd4f0cdc4f7798d0108a2c9f890e114ac179db12b40

Observation ed1dd05c-ad83-499f-881b-79ee669eb433 · outbound

This paper cites The Llama 3 Herd of Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.596218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.596218Z digest=sha256:b2feaa853901860833cad7a1fd5d1a3edffda336a48baf8e71976f9e98fcf0f2

Observation 45d9ce59-dda3-43cc-b9af-960c5bcf32b1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.606852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.606852Z digest=sha256:2f07bbc49efc0f663e70ecb6a59de9960043d673f42317528301e2e47ddb8522

Observation edf07ee9-c9df-4a15-a4be-6b583c55644a · outbound

This paper cites Neuroplasticity and rehabilitation.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Neuroplasticity and rehabilitation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.070739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.615498Z digest=sha256:c5fcf91c87bc2c485665d8c880788df385dd3bd8cd71d592669f34e47eb9fa40

Observation 02aadcec-4134-4e6f-8cd2-42125ae232c6 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.623148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.623148Z digest=sha256:07f5c7b1df7277bff67f1bfc9e45203bbf89d876476bc06f272328b77522be3c

Observation 0f880a2d-6315-4c62-ae82-8b20b54c5721 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.628996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.628996Z digest=sha256:09a90320930d492a873e48110daf6daadf58873d3990510737fa8e51966e3067

Observation a6169206-4609-4bd3-bbd1-aec2b95a5de2 · outbound

This paper cites Cogagent: A visual language model for gui agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Cogagent: A visual language model for gui agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.635778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.635778Z digest=sha256:fa06769e15abdf9154a0566a8bd1d1d348df109cc873fa49c0cc7db559b29ff5

Observation 487972e9-9ee5-409a-80f6-57864f84b357 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Lora: Low-rank adaptation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.642941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.642941Z digest=sha256:a819283cd08f82723eab46ea9d042c48277b5658812e644af47ad777fea44e3a

Observation 831f7ad4-deb4-4d3f-83a0-04b04207cbdf · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.648306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.648306Z digest=sha256:f5d8de6fdc466b0f4ac1480db8f7deec8aa6e2bc104d53e507918ee8480491b4

Observation ed5549b3-3d54-4fb2-9815-d06a63b473b4 · outbound

This paper cites Os agents: A survey on mllm-based agents for general computing devices use, 2024 b.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Os agents: A survey on mllm-based agents for general computing devices use, 2024 b

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:53.014453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.653801Z digest=sha256:7fe93ee26ec2ecc2985b6cdc90776f61fa480cf83e85d78c5df579741cf4185e

Observation 5d76df08-9209-4cf1-9509-330a6d06bb1f · outbound

This paper cites Mechanisms of motor learning in the cerebellum.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Mechanisms of motor learning in the cerebellum

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.990888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.659738Z digest=sha256:e48e0b79d4042e24b8db186e055a84feba50dfe343261705eef625c710298e91

Observation e50f27c3-bb7f-4721-a431-21dfb23fabf9 · outbound

This paper cites Autowebglm: A large language model-based web navigating agent.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Autowebglm: A large language model-based web navigating agent

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.972155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.666070Z digest=sha256:6131f028c522a63f70b29137abb0c82fa0fb3321a18d550cac018916635a5c0c

Observation 96710486-0104-4b5f-8a91-a545d55cb11b · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.672939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.672939Z digest=sha256:80c3d159f07308c0a36da32b758741cc1ee6427a30e5572b5fd034abb6639697

Observation 312fc508-ed40-4cf0-b844-da993d00925c · outbound

This paper cites Visual instruction tuning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.950298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.679108Z digest=sha256:83c43033858c461a2623d96edb28da576c87e9f52c8f67c22fa8e3ca31c3d38a

Observation 18e22ced-d723-44c6-9338-3c010a917d43 · outbound

This paper cites BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning BOLAA: Benchmarking and Orchestrating LLM-augmented Autonomous Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.685669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.685669Z digest=sha256:ba11c1c1de2b3ba9a42c3dceeead246b378ca207bcd10c09d48fde58c16e64e9

Observation ecbd4669-5baa-47a4-8372-5898edb01355 · outbound

This paper cites Agentrewardbench: Evaluating automatic evaluations of web agent trajectories.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agentrewardbench: Evaluating automatic evaluations of web agent trajectories

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.694654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.694654Z digest=sha256:01ef6e922228b6304ff5734459020362dadf5d4c032ed5437a99e488bd63887d

Observation bb49802b-a5cc-49ca-85ed-62f31cdaddfb · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.700711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.700711Z digest=sha256:3b7b8d640f2bb36c8892b750cb57db9b625574feeaf4a9ef5d56a52f1598dea5

Observation a30da84f-5130-42c3-b14c-0118dbad65fe · outbound

This paper cites Gui agents: A survey.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Gui agents: A survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.708341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.708341Z digest=sha256:854571f82a2ee28ddd0c0e43973219c659ccc7d5ddb9fba2d765c9e86126f3cb

Observation 71692cd5-27e2-494b-9579-7bbd73cebc04 · outbound

This paper cites GPT-4 Technical Report.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning GPT-4 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.713835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.713835Z digest=sha256:92c97db9d4fa01f3c35bd2159dedf785348dbd83519ee458fd0a1f94af89321d

Observation bdb4c472-556e-4dc6-9bc3-86bfd27ddd69 · outbound

This paper cites Operator.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Operator

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.918205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.722716Z digest=sha256:f7d62189884fd7641bcde8cdd39bafc19a3f004e26bfdfd1646e1a9123458ee2

Observation 01bd980b-836f-4962-949d-b70a2bec033d · outbound

This paper cites Training language models to follow instructions with human feedback.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.730369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.730369Z digest=sha256:7cfe179dae513f3397572c3c241b0b14d5192f39fe96e1fa6900b476f62e28fe

Observation 24c293cd-f165-4b49-a0e8-e7ca369e0c89 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Autonomous Evaluation and Refinement of Digital Agents

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.737187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.737187Z digest=sha256:81c5425c5e7b01954d31a7f9bdf5a779edc5a726f8e5e119d3a069c60f3055e8

Observation 057d2df6-59df-499e-a5bc-f172508ff7c4 · outbound

This paper cites Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.752711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.752711Z digest=sha256:d4181307711740b33c4168639aca2efe30bcc4923c5a0d0fd6afedbd2faf0ad7

Observation ceb3ec28-6bb5-463b-97f4-03a7fcc6ba6e · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.759166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.759166Z digest=sha256:863542aaaf9ced400f430db9ba481395e4d6e52a50038304d814ba689eb9719d

Observation a3a602f7-27ae-4246-b411-6d38d2dd4149 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.765656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.765656Z digest=sha256:e2d3ac63a19508f12a65e8ad501a099a57b35dfc87e1a6af8b38d74b5eca8092

Observation 032941e2-7c06-4749-97ec-d4baeeb78e46 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.774012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.774012Z digest=sha256:6645cafdf695442e2bdd51b445d735064c40d10f07c1de02f1232ad672cb7cad

Observation 2958264b-f2eb-46b2-8a74-dc358820b2c0 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.779222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.779222Z digest=sha256:8e964b4427868db37fd7a8669c14ae2831f528ffbb869e90f49ce3c4993ab5b9

Observation 68d00cff-0bf1-496a-b6dc-07745af948d0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.785368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.785368Z digest=sha256:8c0889e97cd8d24e53c39aca5752a5a83eb390ab912c49e55acbf71037f3837c

Observation 4cbdfc24-f676-4341-af48-f3f580f3ad50 · outbound

This paper cites Coact-1: Computer-using agents with coding as actions.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Coact-1: Computer-using agents with coding as actions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.794782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.794782Z digest=sha256:1bfb1ea4b0640b44b6183857f6068bc3c3db1bb0d3fce24553fd684b967d5a64

Observation a5e2c418-4155-4d6a-9d03-fa507e6e8482 · outbound

This paper cites A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.801652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.801652Z digest=sha256:7304d527befd3f1b7b42ed8cb6edf2d40009e79a36ef2cc50e3501d7778745fc

Observation 329bea5a-7f2f-4714-adc0-154d95ba13d9 · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.807916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.807916Z digest=sha256:d6a9fb8e147a6255cd11c7ee22df6a7993aed5e9a9167dbe11cdaf3d4536d25c

Observation 40e44535-cd47-4b76-8e5a-ee9bbb0a8970 · outbound

This paper cites ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.814391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.814391Z digest=sha256:98f58828b97c89a54cdf2e261cd9e63b0ffc2a916a99d1e672675ddcade7bed3

Observation 84696db2-b6ed-48cf-9186-844744fabb5a · outbound

This paper cites X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:19:51.797076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.820124Z digest=sha256:f930bb362ca64fe16d4dd35adf6d42491427b41690aecdcaaf8f5eff54c88c7e

Observation aedb1683-460c-42d4-b869-397baa43815d · outbound

This paper cites Bootstrap3d: Improving 3d content creation with synthetic data.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Bootstrap3d: Improving 3d content creation with synthetic data

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.861294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.825509Z digest=sha256:cf001c9a914581d7f2e5e7ca8d440c57e7c0fec630c49fecb7de9bc61e375fd1

Observation 49b4daea-61b9-4c3d-bbb7-96c1f24217f2 · outbound

This paper cites SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.834935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.834935Z digest=sha256:0e437133059578d7c67b0740a6e4ab7bffb4cd92fee4669c9f5aff84ded28423

Observation da93040c-03ae-4654-8271-0c5ef21148b2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.840176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.840176Z digest=sha256:d6fd5ef5628bb5c7479a9428888fdc893fe766d9d1602c87ae9bca559b3adb7a

Observation c8f2b661-b8cf-44b7-8ffc-53fecc639b54 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.845468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.845468Z digest=sha256:ba4cdd5e1a98295fdf29b0e8e12a264cf6f6e8006e8c779ab13685695300d96c

Observation 19680e7c-890b-4461-a2ff-a5fb8342b518 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.851707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.851707Z digest=sha256:65e1c881c1ada1d826f982684da3306c0a685de0813b71d365b63dbab5239425

Observation df0f4c89-4637-4436-b04f-92e8e56b571b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.858856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.858856Z digest=sha256:6a2f2d09f8828b7d02a512c1539d8da93976d9ea178530b5795daeecd56df48d

Observation 797d43fd-7197-41e7-bf5e-968890cc8ce2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.864786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.864786Z digest=sha256:a3a26f9f8555c684703c20cb5a34c0d1919680f9069a7dcfbee83b596a488790

Observation 9907cfa2-4638-489e-95d6-861cf9afaee7 · outbound

This paper cites GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.871679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.871679Z digest=sha256:88fdd031ec98688adfb1b7ef0b056fb0b9eecb4c9f17f6857fecd8d23d39e9cd

Observation ff81c056-6cc3-41b8-be74-c0357b4352f1 · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.877434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.877434Z digest=sha256:6277a28d46cd61119c1af4c60a3be8627ff15682c8b7f7b7f787e8f3734a6bd0

Observation cd5029d8-518e-4fcb-b74f-abc604e30118 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.883825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.883825Z digest=sha256:33926f421ac5d25a6ec27923519c64e2fe1ed2e601590042b5597ff83ad5942b

Observation 44df1c76-9e65-4a08-b564-04d25ab53a75 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.803037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.890161Z digest=sha256:6ac49fbed202415067784c3044b21b66eccc22decc2d39505c72a59b83c3de28

Observation 1e0ef00a-f26e-4e5a-b5f2-e24d4edeb7f1 · outbound

This paper cites Scaling computer-use grounding via user interface decomposition and synthesis.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Scaling computer-use grounding via user interface decomposition and synthesis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.897174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.897174Z digest=sha256:294679c4ba7e0daac9ba163956bc48c29aca457e869a4912924cc4b33d5bad88

Observation 909b480f-cc88-4ceb-92a9-a48cf0da6589 · outbound

This paper cites ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.902512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.902512Z digest=sha256:e385a7271f6d8dc64976ac42a7083ab3a42193c8bbc8b197f206c80372aea81c

Observation 5157dc22-57cf-42a6-9430-acbbbc010f77 · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.914458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.914458Z digest=sha256:064f447e33f1704e9e7a2969c6a0305639144ef6ccb4e6d6adb9e9ff8e9b9abf

Observation b838e4cb-25af-4756-9ced-7248d9a4dedf · outbound

This paper cites Fine-tuning large vision-language models as decision-making agents via reinforcement learning.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Fine-tuning large vision-language models as decision-making agents via reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.775523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.921108Z digest=sha256:d556e3346412054955f6917f7a1808d91bd6dc58a400de7d4c5903c09b044b26

Observation 2ce47284-8dae-4e6f-a4d4-2670257e4e41 · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Appagent: Multimodal agents as smartphone users

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:19:52.753496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-05T15:19:50.928045Z digest=sha256:b0a428d6f5a74ae3268f6ae96cc581d392992ddbd73b263f9cbcd4236c67e869

Observation a6a42c0f-65c9-4b3d-b61c-f0b09b075a31 · outbound

This paper cites Android in the Zoo: Chain-of-Action-Thought for GUI Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Android in the Zoo: Chain-of-Action-Thought for GUI Agents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.933323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.933323Z digest=sha256:9b885098e9e5010112bbae71fe64b04216f2bd3494c53ee064532a34e6c3d631

Observation 62386c80-f36b-4f40-9ee2-524a2c2aa569 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.939908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.939908Z digest=sha256:59a5dc6a3c2cbfe536e67fed2dbd98e2df3a58ddd6647ecacd5fe81106e964bf

Observation 535dcb56-38ea-4662-8ec2-dd8827cf14c3 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.947042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.947042Z digest=sha256:d7bde58cbae9868e415e56774fad31cd6ec84a28b526279c6aaf4964c6c2c64a

Observation 7aa67b8e-3163-4fd6-bab5-b1873660a265 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.953787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.953787Z digest=sha256:403a875dc9864ff4b824affe43652c94a0cfebae4b293bbdd0f60a986ed3e531

Observation e39bfaa4-d417-4291-86f9-0218b6453e8f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.959760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.959760Z digest=sha256:b3944f42ddcbfefd638a65ee15bc6695178d71eec0893c9ea2aea9adbb9a4971

Observation 38ef65cc-3db8-4ecf-ac21-78bee8854e5b · outbound

This paper cites write newline.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.966719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.966719Z digest=sha256:249879e207826015fdccf1e1f1b7c089a0792e349a739701877c7910a96b6aff

Observation 90d3248b-ada5-4690-9aa5-29236f09ef3a · outbound

This paper cites @esa (Ref.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning @esa (Ref

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.972838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.972838Z digest=sha256:88876f764c82565b673d154d42bd731b61874a38d00ba5496186ab889386b44b

Observation 74f7dfa6-d1ef-4e47-bc5d-adf7b091d966 · outbound

This paper cites an unresolved cited work.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.978224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.978224Z digest=sha256:715db12b199315ba76f86c9a45e5428b02024fbc49355425636614dabdd07e2c

Observation 9ecf09ff-f65d-4172-ad5d-0ed4431b759a · outbound

This paper cites planner" and an.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning planner" and an

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-08-05T15:19:50.983692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.983692Z digest=sha256:2fe252aa81450bc8fd60810595379a479fad03e205c14eb017a6a2a6318613de

Pith citing papers

Observation 460e88b7-e6db-4b59-946e-d4d804092b43 · inbound

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents cites this paper.

Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.929498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T14:15:55.180284Z digest=sha256:eb5e3d35f7a8510a0cc51ca26ffe63041f5de22981a95644c9eca51b9477784c

Observation 9d79eac9-d5b5-4c6e-bdd7-ca9363ed2204 · inbound

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models cites this paper.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.817666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.817666Z digest=sha256:0bec56767793f081ab0b420487274f7b3827d1110535b9d6793990830c028382