Pith. sign in

Paper Citation Record · LEDGER

DPO Learning with LLMs-Judge Signal for Computer Use Agents

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2506.03095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03095 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:11:48.561600Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84a62ef0-d64b-4099-8946-7d03bf5faaa2 · outbound

This paper cites Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:46.803506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:46.803506Z digest=sha256:2acc94702dc64b1cdfa5fa274f2d7b84e49921927af1936b8fd4f1acdb3635d4

Observation ea1053d5-30ed-4248-9cb2-640ec29e577f · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:46.892805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:46.892805Z digest=sha256:65d75d925dcdf94b45646d59bb97dd11ec9acf69dbdee9fe0250b70f890cf71b

Observation 05cea6c6-4d04-40f8-9aa7-ee208e20f0b7 · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:46.979114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:46.979114Z digest=sha256:2346c82823dad437d4fcd511c286b849c21d8598a9a3a7ccc1c6bbe22f6787ae

Observation 7d4bee5d-5850-4a70-884f-3bf4e00f7f47 · outbound

This paper cites Deep reinforcement learn- ing from human preferences.Advances in neural information processing systems, 30, 2017.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Deep reinforcement learn- ing from human preferences.Advances in neural information processing systems, 30, 2017

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:51.007382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.046229Z digest=sha256:a55e690ca1add65be667910cfb8b59b75bce4b3b1b33f4635e34445c011db92a

Observation 658122ca-daf3-4bc7-a0ad-804fd2eb11c1 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Mind2web: Towards a generalist agent for the web

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:50.797671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.127378Z digest=sha256:96321b8aa057432a2904760f486328e832f5365bcfc90bca173090d756292386

Observation e9c7f30d-34c8-4991-8b0f-56f0c395b6bc · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Detecting and preventing hallucinations in large vision language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:50.533904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.230549Z digest=sha256:d0856744c00478557dcdee7425d5542d06121616911340e8c7c3c5b0c7a7de8f

Observation 0a2d819d-b698-43dc-bfaa-8625d3a61fd5 · outbound

This paper cites From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge.

DPO Learning with LLMs-Judge Signal for Computer Use Agents From gener- ation to judgment: Opportunities and challenges of llm-as-a- judge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.315267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.315267Z digest=sha256:72bf1e114d79288a0a0a1b48e2701c058a9604643eefaf1274c377e93b360f9e

Observation 2851502b-aa0c-47f8-bb4b-5bbecad3798b · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Silkie: Preference Distillation for Large Visual Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.407418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.407418Z digest=sha256:7e6a497e5ea719ec87fcec5c1f125dac45d1c4a72ba7385a28cbf05d10dfedd1

Observation 6492f508-04b4-4ca8-9b5f-0aecedd78d40 · outbound

This paper cites Visual instruction tuning.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Visual instruction tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.487075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.487075Z digest=sha256:71f5e570d21578c6b19a7f0fbafa8b56f20688b689632af6416159f60340d6dc

Observation 3d6fae3a-0208-4387-b14c-424fea65d435 · outbound

This paper cites InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection.

DPO Learning with LLMs-Judge Signal for Computer Use Agents InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.579318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.579318Z digest=sha256:b5ca765cc2d14267d52c4ae7649245e168c5a8ddb8251c14f387e301e711b491

Observation b84ccf14-510c-4c8b-894c-088bc0000284 · outbound

This paper cites Hello gpt-4o.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Hello gpt-4o

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:50.238584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.641146Z digest=sha256:a976b23dca10f9bebfdd6379fd211d5c5ae0139214c73995a19aa2c8331d86cc

Observation d39b2950-ae32-43f9-99d1-9a49b6d4d6cd · outbound

This paper cites Training language models to follow instructions with human feedback.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Training language models to follow instructions with human feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.985964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.702656Z digest=sha256:82a106757259155b8d00f7c7530d46fce163fb7ab145cf6ffad9817931568de8

Observation 4d16142b-d8d2-4b88-b24d-c4fd0f912613 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

DPO Learning with LLMs-Judge Signal for Computer Use Agents UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.763545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.763545Z digest=sha256:ceb54d950a0776da4a4a57659ba3d64bec70bee32d770bc653ce01870f4d8de4

Observation bef87bf5-1916-4d8e-bc38-e68ae7ee8514 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.695754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.819248Z digest=sha256:06a6470dea9cb42c94158b934c2278b4a18fab35793538e2e98d1111030716ad

Observation c0f62f9f-6d31-4201-bc4e-345989574692 · outbound

This paper cites Androidinthewild: A large- scale dataset for android device control.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Androidinthewild: A large- scale dataset for android device control

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.492838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:47.869310Z digest=sha256:409564e4c6a3fb4ba045917df40ffbd75ad95c6b46f998968fb12f6a4d3bc7a0

Observation 85780b80-96e0-4ffa-8e75-495e650e76ea · outbound

This paper cites Proximal Policy Optimization Algorithms.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.932567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.932567Z digest=sha256:18582da6415125a35a24e3a1c31959c2a7cf4b9ac00029c5b63d67623153829c

Observation 52fc45d8-d8ae-47c0-8146-e8dab32c6959 · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

DPO Learning with LLMs-Judge Signal for Computer Use Agents mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:47.992971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:47.992971Z digest=sha256:3770a1fc22575add86539574d9b1bffbb2fdb891048bb681d4a06870b9483fb9

Observation 13a7da60-6ed7-43cf-b502-223ebe2c8838 · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

DPO Learning with LLMs-Judge Signal for Computer Use Agents OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.072942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.072942Z digest=sha256:080d4ce267302562663a48c2c147867d006fdb3144ea5af00f84dae50e8f04f2

Observation 2c4b1f6a-fc90-4e32-b79d-28036744c2b8 · outbound

This paper cites Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Osworld: Benchmark- ing multimodal agents for open-ended tasks in real computer environments

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.294060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:48.154432Z digest=sha256:5640edd67bf96d8fc60c8a71ed50bc1dc8f990fde993544cdd51265cebb94b03

Observation b940e2a0-ab9d-487c-ae4c-5a991da57c58 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.195477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.195477Z digest=sha256:46c778096da0f5bc0230e9fba06dfcbc6bbcba521cf669f123c8d526e4f84f3c

Observation f3e53efa-5ea7-40c2-a989-a6f99ef929aa · outbound

This paper cites GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation.

DPO Learning with LLMs-Judge Signal for Computer Use Agents GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.250484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.250484Z digest=sha256:bbc056070921a2d281b10e599f523b4ba02a4f534517843e5143b3e03e351d73

Observation a6d1d13c-c917-42b2-b910-fd47e5cd3e2d · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.321722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.321722Z digest=sha256:88e5af7152a8213ff4d45d77e5c9a731378e246da6e844238bc0a46ee095e8d1

Observation 2e782ce6-1f08-47c5-ab9b-cc56e749039c · outbound

This paper cites Gpt-4v (ision) is a generalist web agent, if grounded.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Gpt-4v (ision) is a generalist web agent, if grounded

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:49.083792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:48.389540Z digest=sha256:e3b38f39bb949a5125fc2af4b79dcb1c6b985eb6a3ef0a928fa0b45678295bf7

Observation 2a5c22b9-21d8-49cf-93b3-dda022dd0684 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:11:48.925875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:11:48.452167Z digest=sha256:ecc2f79a76a684e88728fefac280a874a6bff413a6bb1fa1a0997614d0d14871

Observation b83d7520-3411-4e9f-a70c-3042e5b5d92c · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

DPO Learning with LLMs-Judge Signal for Computer Use Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.502654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.502654Z digest=sha256:ae6fb2b96e00bdc72d003abec3f825b517ae11444f9d5ed74c3a5e8b32a1ce6d

Observation 552f0184-948a-4b4d-b4b4-5ce9554e9345 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

DPO Learning with LLMs-Judge Signal for Computer Use Agents Fine-Tuning Language Models from Human Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:48.561600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:48.561600Z digest=sha256:e512392bed49d9d1a42a8de2e47b3d7d0ac466c55f9b08e67e96e914f8e91e83

Pith citing papers

No inbound Pith citation observations are available.