Pith. sign in

Paper Citation Record · LEDGER

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

As of 10 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2506.08933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08933 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:02:29.300721Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T11:13:48.529796Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T11:18:13.822352Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f8405fa-a978-4069-8133-d9e233067f62 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.127878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.127878Z digest=sha256:bb24a228e124677d748ac23a12335cf263b65d12af2b6e20325769865c6a693d

Observation 5c44b633-0590-4036-a2e2-3fd8e01b96ee · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:34.084103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.132482Z digest=sha256:b9fd57dff61448e38b56d52fa17daa290a099d5906b1de3d40847af73ea252f2

Observation 1e40ed52-f436-4e95-88cd-94d4f9cc67fb · outbound

This paper cites Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.136593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.136593Z digest=sha256:52f75e6d6b5d8a2841159323fc324814d01d6229349bff72f68793254e371c82

Observation c585176a-3b37-4fc7-ad61-f9bb97bedd82 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.919293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.140682Z digest=sha256:4b80e75275efc2594bea00e35365e1cdba8bd8678c54a4a8e6e56b9390362dfb

Observation e72355a2-aa45-4aa6-83bc-121b7dbb4d8f · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.144415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.144415Z digest=sha256:affda3904eabca483cf4659452601c2962613a893cfa4f0350c9c6407d105899

Observation 9d9c051e-a413-4673-9b0f-7ebe379f7d5b · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mind2web: Towards a generalist agent for the web

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.734383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.148148Z digest=sha256:76ddb42897297aa59956932897d3c61e7dc8f968a80dc8f7841f9f0cb5eb8c8a

Observation 67c688ee-3875-47df-a4da-3c9fce6c7863 · outbound

This paper cites Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Dysen-vdm: Empowering dynamics-aware text-to-video diffusion with llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.528030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.152268Z digest=sha256:b226b69de6f830d00ce066391f3dd3634f51806df8f1a84151998e4beb272da7

Observation 50aed938-2aa1-4eab-a811-b90eb574a63f · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.349259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.155750Z digest=sha256:c4fcab0967de1c87bde85cf3a2625b14ec05706ff9a3c6033e1ff35a462c7dc7

Observation 1bf59cfd-e644-4886-8d98-1efac1ee8971 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:33.167945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.159175Z digest=sha256:006e859a008aeddbb80a3b03907069eaaca92428dc12d128cef3c16f1d982461

Observation 5cc1e063-7d1b-4e3b-8cfe-2036738d51d2 · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Enhancing video-language representations with structural spatio-temporal alignment

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.995079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.163035Z digest=sha256:1eb926b5ae1ed6ed8126a509ede0a45c964215f7ceee8cae2aaf4d887b2a7fee

Observation 0fcf5a49-4059-4727-be64-f8e5bbd3ba15 · outbound

This paper cites Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.166735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.166735Z digest=sha256:3e1398b76661bc3503ff9583ba648734e2bcb1320dabe0a2d7a121f87d3682ba

Observation b3fe1d08-9cfd-4472-b003-c31dc1ce5fa1 · outbound

This paper cites De-fine: Decomposing and refining visual programs with auto-feedback.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities De-fine: Decomposing and refining visual programs with auto-feedback

Reference 12

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:30.062241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.170447Z digest=sha256:3cbdbb1ccf8bb2c91f169940bbbc3f562278629c3dc438b06842e033646abf20

Observation 37941295-095e-4a19-ba50-e32226ef34bb · outbound

This paper cites Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.173944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.173944Z digest=sha256:ef554b50fb8e04cd7fecc5583dab806c2a3dfb9ab1f5d0f698ba76e028b7fbd1

Observation 47c6f886-57fb-4877-b41b-63b287dd429e · outbound

This paper cites Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.178376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.178376Z digest=sha256:627eeb12143888edae61bc1aeb7b88f62f2c9c85ecc8c7031d7310b83687f413

Observation fe01b38f-2a86-416b-8679-b17ba9bfccf1 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.182555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.182555Z digest=sha256:43ecf69b6bcfee7ac74ca5c0a11d9f348c7801eb8d53227685bbb99a7a832402

Observation 4b378d0a-8832-4d3f-9430-80b301def70a · outbound

This paper cites Cogagent: A visual language model for gui agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Cogagent: A visual language model for gui agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.731848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.186253Z digest=sha256:e076006b79c0a138b4b25152e7d9aebf48b001e8236b0221e5c027437a0fe876

Observation 6eeeba8a-e01f-4a5f-962a-9c4bf82099da · outbound

This paper cites The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.189552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.189552Z digest=sha256:4fcf8b49436f0229945fb82da25b9ffe0c90f080792b9930fd6d2d1691d3fea9

Observation dc06c457-3cfb-4c4f-a88c-688f5b10b629 · outbound

This paper cites GPT-4o System Card.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.193533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.193533Z digest=sha256:0e39e84bcfd1c7fac217d50260da85c54ae0452afdc07d3f5cac84c502cebd2d

Observation 4e48966b-d59f-4071-9c24-86c6d746d836 · outbound

This paper cites P., Russak, M., Koh, J.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities P., Russak, M., Koh, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:32.475848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.197603Z digest=sha256:05ead007cf4bc1b272266a8838c803a170a5d91611cf54691969523dee9761d1

Observation f121153e-fd89-45f5-93ad-6929d8b236cc · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.201086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.201086Z digest=sha256:d61e5f17cef0d28ef02fbf177a52a12fd423578063a6d86f47ef5976bf5e8daf

Observation 09b81ed4-503d-4aec-9e09-16341fd9932b · outbound

This paper cites an unresolved cited work.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:02:32.148264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.204810Z digest=sha256:f7734877ec6649cda897cf11c2c5cdd611b3bd4e20fdcbfa69c3f8e85df962d8

Observation 876b291c-883c-4483-9127-a323e8e9a1f2 · outbound

This paper cites Fine-grained semantically aligned vision-language pre-training.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-grained semantically aligned vision-language pre-training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.782608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.208063Z digest=sha256:80d933e45720120c6756f819c0088b6b735a35147b62299d6bf2f9d36f6cdab4

Observation 00160e87-6b8f-499d-9ed5-1656767e577f · outbound

This paper cites Fine-tuning multimodal llms to follow zero-shot demonstrative instructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Fine-tuning multimodal llms to follow zero-shot demonstrative instructions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.509796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.211153Z digest=sha256:29c9171e619ec9d54436514d264ea866f4f096ca5e99079ec3b790bd45a84a06

Observation 3f872814-3c4a-4412-8a69-c117157f7ba5 · outbound

This paper cites Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Variational cross-graph reasoning and adaptive structured semantics learning for compositional temporal grounding

Reference 24

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:02:29.764865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.214658Z digest=sha256:bfc2da0aa1fef02e7b833051dede1038a414883b7fc1ce6f26a7c3f261a195a2

Observation 0347396e-3fa2-4774-9a3d-65449bb06670 · outbound

This paper cites Mapping Natural Language Instructions to Mobile UI Action Sequences.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Mapping Natural Language Instructions to Mobile UI Action Sequences

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.217639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.217639Z digest=sha256:8c5ed9f9647a4200c76fe61ce8681a543404c4fdad5a464272e692e4c7eca12b

Observation 83b75b66-3873-4d89-8e09-65c9ab7b1b8b · outbound

This paper cites ShowUI: One Vision-Language-Action Model for GUI Visual Agent.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ShowUI: One Vision-Language-Action Model for GUI Visual Agent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.221294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.221294Z digest=sha256:a1f038797dd10f531ac52e4ca257a6481f51134fd3a21a8dbed4a0d5be7f5b47

Observation 089d71e4-e170-48f0-81f2-235ddf624bfd · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.228691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.228691Z digest=sha256:bc7d263fab0183b231532443a5fc69471a7c6811236eea5e0b52fa73aeeb0d23

Observation ae985b10-a123-412a-a6b7-b4d7fac1a85a · outbound

This paper cites Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Boosting Virtual Agent Learning and Reasoning: A Step-Wise, Multi-Dimensional, and Generalist Reward Model with Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.232405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.232405Z digest=sha256:a3a7256172bad4a69f76db4be0446c285979f53b2950cfba3c8bf59c9130f472

Observation 35383744-47a4-4262-8a60-eea83d4ecbb3 · outbound

This paper cites Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Self-supervised Meta-Prompt Learning with Meta-Gradient Regularization for Few-shot Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.236247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.236247Z digest=sha256:7615892747a6cc645c79964f29e48ff5fd1b3ce6fd94fe5ed29fb73fa887ee11

Observation 22983b81-c386-4972-8cf2-d3a86314cd39 · outbound

This paper cites Towards unified multimodal editing with enhanced knowledge collaboration.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Towards unified multimodal editing with enhanced knowledge collaboration

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:31.207365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.239831Z digest=sha256:5f9f4de4a8813b0faf98c37c8083b454960f81a4d79a41da1c9cd7f43740d3b2

Observation 84353a44-ec40-4be6-9285-b5c2ae7fab47 · outbound

This paper cites I3: I ntent-i ntrospective retrieval conditioned on i nstructions.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities I3: I ntent-i ntrospective retrieval conditioned on i nstructions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.908383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.243088Z digest=sha256:80a1b89422e8c9653a6cedd27bed98391b524d055968905387627602787563ae

Observation 17e9e6e2-2b33-4e2e-8b5b-99c661c349f5 · outbound

This paper cites Auto-Encoding Morph-Tokens for Multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Auto-Encoding Morph-Tokens for Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.246334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.246334Z digest=sha256:4ec9a0713c0bc0ee4434ee6ae85a0ada3a489737e4f11c4ff64653c4872e5498

Observation 9ae4e593-e82f-407f-bafa-0a8250e4ea4f · outbound

This paper cites Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.249982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.249982Z digest=sha256:075df8c228f1d1510396ed1ec257e3d210cab7a73bbc2f4315ed2e673ce7b1f6

Observation 2d9bacd5-bd51-44a3-be3e-09a722465bc2 · outbound

This paper cites Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Unlocking aha moments via reinforcement learning: Advancing collaborative visual comprehension and generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.254238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.254238Z digest=sha256:169fdf249726e306cd886bdfd74e6dc5d1a13ba0a55559c9190e912a4d677e8d

Observation 4a7d1a39-3a00-4c29-a01b-fa13b471f726 · outbound

This paper cites Androidinthewild: A large-scale dataset for android device control.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Androidinthewild: A large-scale dataset for android device control

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.570748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.258032Z digest=sha256:7019dc0594d9c203e3d60471cdf1e85473b7e6a91808978eda8a617e748539ce

Observation 215b55d8-7e85-4afa-aed2-eeb7a5b6b12d · outbound

This paper cites ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.261946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.261946Z digest=sha256:eb792d80f29b818ba4065ccdb6bab8706b644877606a1595ed3006370e992925

Observation 06084d52-bd1f-4843-ba0b-22d00827e86e · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities TaskBench: Benchmarking Large Language Models for Task Automation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.265475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.265475Z digest=sha256:9882ccaf957341bd40afd407a02bd0f4492eebcb46a2698dad12766f981b9524

Observation dc08f448-c10b-4727-b331-adce39e1ed8a · outbound

This paper cites MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.269282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.269282Z digest=sha256:d41930f81d1c0065579bd18eb4a6f79766ded9f005f1a8e0bb514fb8fcaac5fb

Observation 11c0cfee-b8d0-4148-a46d-5aec8bc47488 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.273032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.273032Z digest=sha256:4b8ae09cd3db4ad7465e578ccebfa79906c991f68f153a84a74b1fb4d5195bca

Observation 68938c66-c337-432b-a812-8693b38d13aa · outbound

This paper cites NE x T - GPT : Any-to-any multimodal LLM.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities NE x T - GPT : Any-to-any multimodal LLM

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:02:30.239250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T05:02:29.276630Z digest=sha256:8c8273bf41c69098772f7f8e8a121830452d4bd4f4e5708b7f64579cbbd95f3e

Observation 08be7f69-11fe-49c2-9877-76c527e51f68 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.280843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.280843Z digest=sha256:c555badeb43fa17afc5fd5704d9fc96c4a558b0e67eb130d75170d7485c93df9

Observation bb747de6-b1f3-453e-8a2f-ccdbc3835f84 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.285064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.285064Z digest=sha256:17a7c250bf2dd8c03dce5d4f9b352c707111823fd24bed385ca16dc0e4f29d3c

Observation c5611f9f-14ea-4cb8-95f4-037ad2065bbb · outbound

This paper cites CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.289598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.289598Z digest=sha256:568024c4f9932a6b319ca41915928054c597259ecde90d250ce6fecb78ba256f

Observation 9f40b8c8-a93a-4918-a518-b39157dca0fd · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.293289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.293289Z digest=sha256:f7abd1816392f4b535fce17795a09af1afc7d766bad7ab0b1a265a57d07aa53d

Observation 6febf858-5e25-45ba-986e-e62d8cbc97f9 · outbound

This paper cites Webshop: Towards scalable real-world web interaction with grounded language agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Webshop: Towards scalable real-world web interaction with grounded language agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.297322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.297322Z digest=sha256:abdc45e8166db275988af9dd3471a8bd2366c73524c9bed091f568caab418046

Observation fd8ab3c2-cc57-474f-b695-c15d6a274724 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.300721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.300721Z digest=sha256:5b06194cabe5934fe1647e254b3a64da32f64bec97410a0818d5dbbdc35b3931

Pith citing papers

Observation 337ed338-7df9-4130-b2f5-f60bf7e91408 · inbound

DocOS: Towards Proactive Document-Guided Actions in GUI Agents cites this paper.

DocOS: Towards Proactive Document-Guided Actions in GUI Agents What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.823868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T11:13:48.529796Z digest=sha256:60e4ff94b66b225b49883dfba002064082e275e94d383c6ad5fa5ab18fead795