Pith. sign in

Paper Citation Record · LEDGER

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

As of 17 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2506.08933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08933 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:02:29.300721Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:54:51.906676Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:18:13.822352Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f8405fa-a978-4069-8133-d9e233067f62 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.127878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.127878Z digest=sha256:61c51c2942bfd3eb80d84538443163d4dd610bfb26f6de39f80b8348fb754acd

Observation 5c44b633-0590-4036-a2e2-3fd8e01b96ee · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:34.084103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.132482Z digest=sha256:2e9c0130a6cdb471d937e68261e28aa3a95b21a4b909e32b4e67abf9c086fcb4

Observation 1e40ed52-f436-4e95-88cd-94d4f9cc67fb · outbound

This paper cites Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.136593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.136593Z digest=sha256:bbb704801f5489138d5afbeb4e22f3e48a95f4f694a7687c4602fb32c59c08d7

Observation c585176a-3b37-4fc7-ad61-f9bb97bedd82 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.919293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.140682Z digest=sha256:a7c7283c3a804210b01f268273f092c11a86a2da0dae712df97088910301e05d

Observation e72355a2-aa45-4aa6-83bc-121b7dbb4d8f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.144415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.144415Z digest=sha256:8860141a1222c7df41ebf6b3db04c181abdce77867cc285daa517b8fd1a41403

Observation 9d9c051e-a413-4673-9b0f-7ebe379f7d5b · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mind2web: Towards a generalist agent for the web

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.734383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.148148Z digest=sha256:38dbdceff17185bd0f6f98de00a00d65fa437a23f77a7e9bb5d41721f6106998

Observation 67c688ee-3875-47df-a4da-3c9fce6c7863 · outbound

This paper cites Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.528030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.152268Z digest=sha256:6fa65b2a68c40c7f46b9c5cd58117d58118c9686e2288098ac98371e65d13bb2

Observation 50aed938-2aa1-4eab-a811-b90eb574a63f · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.349259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.155750Z digest=sha256:82f9d035b33fa36f48e9bde9c1ee30085b5db0da57cb6fe50eb87401c86a2832

Observation 1bf59cfd-e644-4886-8d98-1efac1ee8971 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.167945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.159175Z digest=sha256:f6e807b4794c9fd6b5c73963e5c404be00ddcf086fa1f8ae9a7ffbbe77bffcb8

Observation 5cc1e063-7d1b-4e3b-8cfe-2036738d51d2 · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Enhancing video-language representations with structural spatio-temporal alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.995079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.163035Z digest=sha256:dbdebcac6e5512e60904d30c99664e5db5a29890cd642d691cd964e425619adb

Observation 0fcf5a49-4059-4727-be64-f8e5bbd3ba15 · outbound

This paper cites Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.166735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.166735Z digest=sha256:6b7d07c9004ce61944263f9b7a0624d7d48feeffc4e0a614069585a8cd3f5feb

Observation b3fe1d08-9cfd-4472-b003-c31dc1ce5fa1 · outbound

This paper cites De-fine: Decomposing and refining visual programs with auto-feedback.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities De-fine: Decomposing and refining visual programs with auto-feedback

Reference 12

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:30.062241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.170447Z digest=sha256:63aca781f53a03c4c71aee88dbf7918ca3a75cf5812163f958156f2159f3342c

Observation 37941295-095e-4a19-ba50-e32226ef34bb · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.173944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.173944Z digest=sha256:b7cfa0c56cca55059110b9012bfbe9db75db9ae7069665d88b6c1cf72e75ff20

Observation 47c6f886-57fb-4877-b41b-63b287dd429e · outbound

This paper cites Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.178376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.178376Z digest=sha256:acc0857a8fcf1b97d0407ee695d4d63bae74705607357dab3b40bc5f6618fdbf

Observation fe01b38f-2a86-416b-8679-b17ba9bfccf1 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.182555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.182555Z digest=sha256:61aea85e87c5331a252a639c9800b436f26e762e1eba4a26e22b3cfedaf1b901

Observation 4b378d0a-8832-4d3f-9430-80b301def70a · outbound

This paper cites Cogagent: A visual language model for gui agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Cogagent: A visual language model for gui agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.731848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.186253Z digest=sha256:30a23eec0a0fdf9e9d4fa143488e55ca350a9a94c544737be65834d933c3d2c8

Observation 6eeeba8a-e01f-4a5f-962a-9c4bf82099da · outbound

This paper cites The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.189552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.189552Z digest=sha256:71e59534c27f615aab914cd7aced758fb7b7e280425cad4e91516f3b2a7f1b79

Observation dc06c457-3cfb-4c4f-a88c-688f5b10b629 · outbound

This paper cites GPT-4o System Card.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.193533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.193533Z digest=sha256:576927b9dc73bea34b604b63e6a93f5cdcdd6b474383bc85c964c57c364d3e74

Observation 4e48966b-d59f-4071-9c24-86c6d746d836 · outbound

This paper cites P., Russak, M., Koh, J.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities P., Russak, M., Koh, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.475848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.197603Z digest=sha256:801437644542e69f9e6604d9f28bd1b9fd5bcd009a45615824d19f45299e59bc

Observation f121153e-fd89-45f5-93ad-6929d8b236cc · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.201086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.201086Z digest=sha256:d63d3f0ecbced2adf24edc6996bdbdbf002af85a97c2f2a25c1f8c54c8e758b2

Observation 09b81ed4-503d-4aec-9e09-16341fd9932b · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:32.148264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.204810Z digest=sha256:23ca334f301990bef61b1ea5d6b2bd997620ea3cea542ca84b48ef3b4573a837

Observation 876b291c-883c-4483-9127-a323e8e9a1f2 · outbound

This paper cites Fine-grained semantically aligned vision-language pre-training.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-grained semantically aligned vision-language pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.782608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.208063Z digest=sha256:14bc60bd5e79c56de894c78f3ee1917f3d39182ca9d8131c3f5d09067e0ffa77

Observation 00160e87-6b8f-499d-9ed5-1656767e577f · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative instructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-tuning multimodal llms to follow zero-shot demonstrative instructions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.509796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.211153Z digest=sha256:cb1ffcfe0f3b43b94042c3ccb729806ca2f325dd7ae2f689265ec9fb75741bdb

Observation 3f872814-3c4a-4412-8a69-c117157f7ba5 · outbound

This paper cites Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:29.764865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.214658Z digest=sha256:6545ef44d31127ddd6bb02f8733def4dcc329c8e7d3a10be7bc21982343bedcd

Observation 0347396e-3fa2-4774-9a3d-65449bb06670 · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.217639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.217639Z digest=sha256:a297abc793f1502b3a04d5f3f2ceb4adb2855d1ea2cf58e8a6b5b0f8801e9832

Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.221294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.221294Z digest=sha256:8b2fd707f880e7a419f200de076ffce9f8ab815e6486e95add41d8122c7d0632

Observation 089d71e4-e170-48f0-81f2-235ddf624bfd · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.228691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.228691Z digest=sha256:b9b7b11bf78e9d481ab94962199931e28f4d61bd190f18265abfa041847080f8

Observation ae985b10-a123-412a-a6b7-b4d7fac1a85a · outbound

This paper cites Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.232405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.232405Z digest=sha256:dc2df3ffe514fc30204ddd6145cd53ab13a437f7fc48692788acb58f513115c5

Observation 35383744-47a4-4262-8a60-eea83d4ecbb3 · outbound

This paper cites Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.236247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.236247Z digest=sha256:bf3efd4db25328bf169388cfbb4eb8afd8ce66e189e9a1fd92082e0c1b87626f

Observation 22983b81-c386-4972-8cf2-d3a86314cd39 · outbound

This paper cites Towards unified multimodal editing with enhanced knowledge collaboration.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Towards unified multimodal editing with enhanced knowledge collaboration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.207365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.239831Z digest=sha256:b1b2259fa6da225b01fc9647f46156bb58985f20f85039df6af672f38e176b12

Observation 84353a44-ec40-4be6-9285-b5c2ae7fab47 · outbound

This paper cites I3: I ntent-i ntrospective retrieval conditioned on i nstructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities I3: I ntent-i ntrospective retrieval conditioned on i nstructions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.908383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.243088Z digest=sha256:4679f56509ccf6c7df728168774656b7129f36b52608dca5e3aee9ce20ae1023

Observation 17e9e6e2-2b33-4e2e-8b5b-99c661c349f5 · outbound

This paper cites Auto-Encoding Morph-Tokens for Multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Auto-Encoding Morph-Tokens for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.246334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.246334Z digest=sha256:a7dc83c0f58ea9894e53cfc9792d8567cb82da5d951127e5ed2bbffb901b365b

Observation 9ae4e593-e82f-407f-bafa-0a8250e4ea4f · outbound

This paper cites Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.249982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.249982Z digest=sha256:0384815423bcdb330e41c0ae5476b9078461ba161e56d8e077d9d5d112e96be4

Observation 2d9bacd5-bd51-44a3-be3e-09a722465bc2 · outbound

This paper cites Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.254238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.254238Z digest=sha256:8398e64ed788c6a3355e98238cade1872504ae7c806a86efafa5b754ec1fad50

Observation 4a7d1a39-3a00-4c29-a01b-fa13b471f726 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Androidinthewild: A large-scale dataset for android device control

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.570748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.258032Z digest=sha256:b218b6c28639847a72b1c0db5ad2b68da7219a55059b77db0a6717892c6b5cde

Observation 215b55d8-7e85-4afa-aed2-eeb7a5b6b12d · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.261946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.261946Z digest=sha256:6e8cc1b1ca380c95819c75a11e51765db749c7d0520390affb5113f13ae1d12a

Observation 06084d52-bd1f-4843-ba0b-22d00827e86e · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities TaskBench: Benchmarking Large Language Models for Task Automation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.265475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.265475Z digest=sha256:c3808f3cffd67791a889ab552ebcb3b9f765ca8b261d2e364844906a12ffe429

Observation dc08f448-c10b-4727-b331-adce39e1ed8a · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.269282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.269282Z digest=sha256:8f11876dbe90a4b75670f575f681e474d036fef82ccdfc6c65a24b29078df1ef

Observation 11c0cfee-b8d0-4148-a46d-5aec8bc47488 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.273032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.273032Z digest=sha256:1d5ba66c364ad7b9dfd7f897029663605323bc45b36d7e5be83a34d7b8cc872c

Observation 68938c66-c337-432b-a812-8693b38d13aa · outbound

This paper cites NE x T - GPT : Any-to-any multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities NE x T - GPT : Any-to-any multimodal LLM

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.239250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.276630Z digest=sha256:85aa0c9aa1fa8c836b914ef8d58a647034841c97b2d35f6da2c02b855a60ce55

Observation 08be7f69-11fe-49c2-9877-76c527e51f68 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.280843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.280843Z digest=sha256:7da128b06d91384c91fbdfeeace52ac84aa05d3dc732093baf8eace96a5fe1cb

Observation bb747de6-b1f3-453e-8a2f-ccdbc3835f84 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.285064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.285064Z digest=sha256:4299da3eef549b3270083b39bc691eeed9a2e630caf879976eee02249c4f7f0f

Observation c5611f9f-14ea-4cb8-95f4-037ad2065bbb · outbound

This paper cites CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.289598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.289598Z digest=sha256:22d00bec539250450722d70d601646bd7603c8091e65754cf60cf59ce357c570

Observation 9f40b8c8-a93a-4918-a518-b39157dca0fd · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.293289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.293289Z digest=sha256:ee9d628e120cd7ef916dbff5e36e1df3ab8bedb6d61108098b60ac41a827757e

Observation 6febf858-5e25-45ba-986e-e62d8cbc97f9 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.297322Z digest=sha256:d01c43b7866023ac623c4e48baeaa989b56115bb3890d13378ee63ad47f743a0

Observation fd8ab3c2-cc57-474f-b695-c15d6a274724 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.300721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.300721Z digest=sha256:a7403629f30b69f013cc675d3d8aa60e0872fe61e2c15b9bd4cea258cea83e91

Pith citing papers

Observation 553be00a-773f-471d-b90d-b411981a531c · inbound

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness cites this paper.

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:54:51.906676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:54:51.906676Z digest=sha256:1b1744162ea457878ac3a513bd9609834d732828409097567e9bea3216aee283

Observation 337ed338-7df9-4130-b2f5-f60bf7e91408 · inbound

DocOS: Towards Proactive Document-Guided Actions in GUI Agents cites this paper.

DocOS: Towards Proactive Document-Guided Actions in GUI Agents What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.823868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T11:13:48.529796Z digest=sha256:ccc4a0a20cd504c20d6e995c5547e32eb1f414018b1dd31bd6e005fb7bf71d79