Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:48.561600Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2506.03095.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:48.561600Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84a62ef0-d64b-4099-8946-7d03bf5faaa2 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea1053d5-30ed-4248-9cb2-640ec29e577f · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cea6c6-4d04-40f8-9aa7-ee208e20f0b7 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4bee5d-5850-4a70-884f-3bf4e00f7f47 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Deep reinforcement learn- ing from human preferences.Advances in neural information processing systems, 30, 2017
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 658122ca-daf3-4bc7-a0ad-804fd2eb11c1 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Mind2web: Towards a generalist agent for the web
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e9c7f30d-34c8-4991-8b0f-56f0c395b6bc · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Detecting and preventing hallucinations in large vision language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a2d819d-b698-43dc-bfaa-8625d3a61fd5 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2851502b-aa0c-47f8-bb4b-5bbecad3798b · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Silkie: Preference Distillation for Large Visual Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6492f508-04b4-4ca8-9b5f-0aecedd78d40 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Visual instruction tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6fae3a-0208-4387-b14c-424fea65d435 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84ccf14-510c-4c8b-894c-088bc0000284 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Hello gpt-4o
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d39b2950-ae32-43f9-99d1-9a49b6d4d6cd · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Training language models to follow instructions with human feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4d16142b-d8d2-4b88-b24d-c4fd0f912613 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef87bf5-1916-4d8e-bc38-e68ae7ee8514 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Direct preference optimization: Your language model is secretly a reward model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0f62f9f-6d31-4201-bc4e-345989574692 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Androidinthewild: A large- scale dataset for android device control
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85780b80-96e0-4ffa-8e75-495e650e76ea · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fc45d8-d8ae-47c0-8146-e8dab32c6959 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a7da60-6ed7-43cf-b502-223ebe2c8838 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents OS-Copilot: Towards Generalist Computer Agents with Self-Improvement
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c4b1f6a-fc90-4e32-b79d-28036744c2b8 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b940e2a0-ab9d-487c-ae4c-5a991da57c58 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3e53efa-5ea7-40c2-a989-a6f99ef929aa · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6d1d13c-c917-42b2-b910-fd47e5cd3e2d · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e782ce6-1f08-47c5-ab9b-cc56e749039c · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Gpt-4v (ision) is a generalist web agent, if grounded
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a5c22b9-21d8-49cf-93b3-dda022dd0684 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b83d7520-3411-4e9f-a70c-3042e5b5d92c · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552f0184-948a-4b4d-b4b4-5ce9554e9345 · outbound
DPO Learning with LLMs-Judge Signal for Computer Use Agents Fine-Tuning Language Models from Human Preferences
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.