Pith. sign in

Paper Citation Record · LEDGER

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

As of 19 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 6 inbound Pith citation observations for arXiv:2412.10342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10342 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:10:29.052921Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:05:09.289637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:08:27.767412Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 462e3955-cf34-447e-927d-7ae20125979e · outbound

This paper cites GPT-4 Technical Report.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.867498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.867498Z digest=sha256:5c271240dd0794d55df29ddeea04afb3baa435957989a9ff6bc824b98965e581

Observation c4ad8e76-eebd-4533-aa3b-301bea25e978 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.626618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.871745Z digest=sha256:184876e5c9b33e5992111ed48048d9739c6b75ef3acba0f5e3786a661ed67be0

Observation e63e462a-5b70-42a4-99ee-39e041a9fc6b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.876033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.876033Z digest=sha256:5288a6d701ba1a97d67a0030d1d9ab48266b5f4d6b72d5e8250ae06062954576

Observation 597c723d-1b13-47e2-9d6d-a7426cb03235 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents, 2023.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Fuyu-8b: A multimodal architecture for ai agents, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.613252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.880833Z digest=sha256:a58606c507c99c6392239bc40ea5ba6a689b1e70a6858423e9ecf3e915cb7ba8

Observation cbce299e-ea7c-4a85-94ba-5118a2da68df · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.886036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.886036Z digest=sha256:bc4b29d141d0238aa8cf7dc3ee878fba30e043e31be2a2faefaaf4b1f2980127

Observation bd8dfdc6-b39b-41ec-9778-7ec9d835ba5e · outbound

This paper cites A computational approach to edge detection.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining A computational approach to edge detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.599787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.891368Z digest=sha256:48533290da5da196bbc876f51bc2443a05eb945884a26ae1bdede3dc1541930b

Observation 28b90062-3e2e-427a-b6ec-1f8a01c43495 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.895775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.895775Z digest=sha256:07b4467810b94a3615c805df2e5d8302855368c855f66c7c6863cd72422a66af

Observation 8ff0e854-320f-4a49-b9da-92e317c5fbe5 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.900333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.900333Z digest=sha256:9739ccbd5100b49411ed7916427e3806fa0199a589bd5d890c01215126fabd0d

Observation f2e8ef3f-266a-4633-810c-aa30d6e6b477 · outbound

This paper cites Rico: A mobile app dataset for building data- driven design applications.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Rico: A mobile app dataset for building data- driven design applications

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.586082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.905456Z digest=sha256:3ce5a3235bbcb4692e192c251e598d58647be139b25b5ac1d0e0dee6a93b8e65

Observation 16f9c842-52c8-460f-980c-c3003bb7ad60 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Mind2web: Towards a generalist agent for the web

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.571928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.910089Z digest=sha256:a61a40089667d5281c457b2b57c7fe06893b2b1d097a144bac8747843a240b57

Observation 9161f461-1e60-4f8f-ad7d-c8fab868559b · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.919795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.919795Z digest=sha256:a1634165f8085122c6edcee10d03c331c02fffa1ce60cdbd2fd8baf112010103

Observation 1ecf13a7-a279-4d13-9cbc-f554050ecf88 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.558706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.924452Z digest=sha256:ac546e3754aa41a18b367affdac71573776f51957b31340a70be24e869af74ea

Observation 773d8398-3400-446b-b473-f82fb8d2d359 · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Enhancing video-language representations with structural spatio-temporal alignment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.544854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.929413Z digest=sha256:c98dc780c55270b1821feadd45a41eb067f3861f9658a2e0661d2c5bce0ee8aa

Observation fb87489d-35a0-4b1b-8583-082cafbd7f0b · outbound

This paper cites Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.934341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.934341Z digest=sha256:4408a08946acf3f84f36293c2964a74170cd85e20cdd68c5e0db1c0e0118be95

Observation 84e459e8-4348-4432-aee6-c3c6fbecdebe · outbound

This paper cites De-fine: De composing and re fin ing visual programs with auto-feedback.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining De-fine: De composing and re fin ing visual programs with auto-feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.529117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.938874Z digest=sha256:712d4398f892a2f85b0f94d0333b6780c75cc0108b3a950fbe54010076c2d85c

Observation 6c3b1587-661a-4697-bd1d-429b24c7400d · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.942871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.942871Z digest=sha256:ee4e03eff0e3c2514321dde62acf9209d83cdc05e90bac20b8ef22fd409f5d5d

Observation 22c75973-7a83-4169-89d2-a9460907a033 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining CogVLM2: Visual Language Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.947119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.947119Z digest=sha256:fb27806cb2975694a30d4deba3aada9ae65b5fe9ee09f9cec53981d69ecabdd5

Observation a3c0777f-3afa-4d7d-b4da-da34b226b0d3 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Cogagent: A visual language model for gui agents

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.515180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.951617Z digest=sha256:675f9b9d9dafad092e47bd6089daaded1d189b3b28720c15e0cb6bf7b0edf36e

Observation 1aa0533d-9155-4c66-a343-9d715d165d60 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.955772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.955772Z digest=sha256:9cca0878fc043ed0e6bc4a53aac52b77ceb4da05537fb1dcde922af8fde7ca4a

Observation d4cab38e-5f1c-4b08-9d62-1bf1787e5236 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Adam: A Method for Stochastic Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.960093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.960093Z digest=sha256:ab3629bce13a3c4417622fed171bd57ec8c408353aa9b4aa3d1b120634e238a1

Observation f944425b-8ed8-4f99-bccc-f632f40237ad · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.964021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.964021Z digest=sha256:ef00dc00e20db127af259ffb3367d548c8c26c937c30f6dfb0bcf436299fc3b5

Observation 129b2dc6-8eab-42b7-a7c6-cc38c7c21a3c · outbound

This paper cites Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.968225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.968225Z digest=sha256:8aed8505c88af4a4799d65a755f3dffb427eab3bb8dc68582bbbc5cd14d2cb4d

Observation 289ae9cd-18e5-408e-a7a0-98ad1fee55ee · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.973202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.973202Z digest=sha256:f067a530ad58d21e36bf6329a9af1c506b26184b46747ec1b42daa6b71ce5ec9

Observation 15d50b90-e9ce-4b7e-a6f2-044a0d703951 · outbound

This paper cites Visual instruction tuning.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.501267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.977592Z digest=sha256:2f3283587205cd43527caab9967beb727a5112bb3fc0a86a5950df8a2d564905

Observation c24df63a-9a42-40c3-b1d7-218e41340dab · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.981478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.981478Z digest=sha256:3969a3a7a17a5f02750ee69a0085affcd2e95f47b2621ecc60d6c81797395cb0

Observation 40a18b4c-7574-4144-be21-f9f3feecea53 · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.985979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.985979Z digest=sha256:cc8c7aec57c235b7634ac5caf1c71d0bd0f17d7bac9814b19464e01d543229c4

Observation 380ec206-46c8-4459-818a-ea52b28e16a6 · outbound

This paper cites Androidinthewild: A large- scale dataset for android device control.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Androidinthewild: A large- scale dataset for android device control

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.487904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.990556Z digest=sha256:96f2effd49ce50a8ff305d458907bac5420c84f56d153e1aa3bcb224caabbbfc

Observation 267af3bb-898f-4b25-a1a4-f66b759c0e5b · outbound

This paper cites META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.994614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.994614Z digest=sha256:775cbecc5804f9099552ac1a5757e1244370cab70d06c44750e27aede410fd6f

Observation ecc04415-d9a4-4a04-b8e2-a5d56b3acf59 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Gemini: A Family of Highly Capable Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.999032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.999032Z digest=sha256:b391f90b9187d035b0a2fe48ebe8bacb2929dcb564d9fa5a383b1e15769443da

Observation a81b2c2d-c420-4e5b-b61a-f9a6127e8a3b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.003225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.003225Z digest=sha256:bcc4c69abc04929d7ca1c0152ec1651b6b431724a20f86836fa3d326143b86c6

Observation 6adb5cd8-dc24-41e8-aa01-80ca931d4788 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.007597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.007597Z digest=sha256:4c40691d81427e590bf0e665532087fdd522bddd510aa0f305d27d48d24384d0

Observation d45563af-ff1d-4516-9a54-3295ee6150cb · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.011807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.011807Z digest=sha256:20ae3d515164ff6764e50bb8c3fcdbb65c8aa4bac93eb025abc870dca536b107

Observation 9a5de42d-6139-49b2-8e43-8918c866aafa · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.016231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.016231Z digest=sha256:9cbf4b81b5a67fa5c944704ec638f46004985806e29ad6fed470595a1078db87

Observation ebddd9d9-b86a-4cc9-ac30-ded666f716c7 · outbound

This paper cites Webui: A dataset for en- hancing visual ui understanding with web semantics.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Webui: A dataset for en- hancing visual ui understanding with web semantics

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.474038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:29.021156Z digest=sha256:cb6ef99bbef4d22e2b0d87c1bc6b9ace78cdc35d6c096b940bb2c6dbe1baa0f7

Observation b2baf02b-11ac-4e2e-a2ef-d21b6122ef8e · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.025490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.025490Z digest=sha256:95a5ad018ecceea5d3cd215f196dcf25833acf659e94b3750096e1b496c7b416

Observation a0472031-8831-4bed-a82c-bec1764fb727 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.030128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.030128Z digest=sha256:9976acffb714d25dd73be1a0124fff4d505d5d6e3d48b3144490094d95fe8e32

Observation c5859b87-0094-4be4-b9a5-d5c06f4380e6 · outbound

This paper cites Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.034981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.034981Z digest=sha256:558657a1d29f72c5e93d0193c2c8c14ce3b47e0a401ac2dd8f9b6bd6622f34a2

Observation 80058f2c-f2a8-43f9-8ec5-975a8065a7b3 · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining UFO: A UI-Focused Agent for Windows OS Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.039295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.039295Z digest=sha256:ab78debda3cafa2cbdd2014d5b5daa12bef208bcd465c4556f58755108f01cdd

Observation 94f898c5-0801-4a50-83b5-d594cfead54d · outbound

This paper cites You Only Look at Screens: Multimodal Chain-of-Action Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining You Only Look at Screens: Multimodal Chain-of-Action Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.043380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.043380Z digest=sha256:7fc4ff495930c9b76c57e4b5a467be3237b80a72f2b00e17fa36a75ce48d2711

Observation 7453521a-0ec9-4fb6-aeb3-7a19c1943f9f · outbound

This paper cites AgentStudio: A Toolkit for Building General Virtual Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining AgentStudio: A Toolkit for Building General Virtual Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.048625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.048625Z digest=sha256:4be66ef1364f87f86cec10566af4c4120461396c0bb1241e61cd792cbcff7a64

Observation ff225322-4c23-4035-a16e-c71e1281f2f0 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.052921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.052921Z digest=sha256:ee880aef53f879d670466f3c6f967b5cb7dafa4b920c7dfa8426bcad9c451ea5

Pith citing papers

Observation f7565c5c-5909-4277-bacb-4540ca8286ba · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.769205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:76d602b554a0a2612a0136393a3536a7d19d54838cca81c7079d45b707403f13

Observation 47c6f886-57fb-4877-b41b-63b287dd429e · inbound

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities cites this paper.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.178376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.178376Z digest=sha256:1c72eda35da2d61e9277d60844a3c3fcd93c6444b517906b974e9af6d0e2602a

Observation 77fc636a-8af5-4362-9d75-3febd07cdabc · inbound

Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills cites this paper.

Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:36.675559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:36.675559Z digest=sha256:e9f8f72e1f27e24e19728840c585ec65005202294fe555eff8546f76f6dbc82e

Observation 3befbff8-0d15-4f3a-8a88-950c2955dd53 · inbound

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? cites this paper.

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:05:09.289637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:05:09.289637Z digest=sha256:e645af21d4f35972c342ab26fb647f6482d7e53913e9d018827562053d7d3fd2

Observation 74b247d5-9142-405e-9f6f-6930299855d2 · inbound

Cybernaut: Towards Reliable Web Automation cites this paper.

Cybernaut: Towards Reliable Web Automation Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:21.052024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:21.052024Z digest=sha256:5472a4fbe533916d072a57565bf0eb8d954268e0622f65f30bbd8fb0bf0137ab

Observation 5c57e7fd-9a60-4885-8155-74ba7bfae981 · inbound

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs cites this paper.

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T16:51:32.578220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:51:32.578220Z digest=sha256:d5605f893c0a71a0876b79125ebae8c89454bbb235f7d5a7a00bd42dd960766e