Pith. sign in

Paper Citation Record · LEDGER

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

As of 18 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 6 inbound Pith citation observations for arXiv:2412.10342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10342 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:10:29.052921Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:05:09.289637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:08:27.767412Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 462e3955-cf34-447e-927d-7ae20125979e · outbound

This paper cites GPT-4 Technical Report.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.867498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.867498Z digest=sha256:6f78147ea0de8400482398727d051c1c8d2f8ba50e26a9720f67da31f0f2dde1

Observation c4ad8e76-eebd-4533-aa3b-301bea25e978 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining The claude 3 model family: Opus, sonnet, haiku

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.626618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.871745Z digest=sha256:8dea282269cf47047ffb810208a4077b16ef43ee1234c6650bbf15e31957441e

Observation e63e462a-5b70-42a4-99ee-39e041a9fc6b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.876033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.876033Z digest=sha256:498a1da27307ab97a26ec4ce0826713942741e02eea6d385e65f45e823368d01

Observation 597c723d-1b13-47e2-9d6d-a7426cb03235 · outbound

This paper cites Fuyu-8b: A multimodal architecture for ai agents, 2023.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Fuyu-8b: A multimodal architecture for ai agents, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.613252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.880833Z digest=sha256:a85240e4cd515b97e119ee9b07d68612ed736cd486595f9f7e00220b5d197874

Observation cbce299e-ea7c-4a85-94ba-5118a2da68df · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.886036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.886036Z digest=sha256:6e5a8e6ac471fe50ec8e0026989925a56fe8baeb3308a86b2aad20724f903b22

Observation bd8dfdc6-b39b-41ec-9778-7ec9d835ba5e · outbound

This paper cites A computational approach to edge detection.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining A computational approach to edge detection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.599787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.891368Z digest=sha256:0c796e11ac80c452ce6e4008f383067495bd181a92dff7f458ce2e4d54b469b6

Observation 28b90062-3e2e-427a-b6ec-1f8a01c43495 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.895775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.895775Z digest=sha256:82bc2ac9b8a0322fa7cc8d306240b386ee95ec9512ae8a914078fd8088201a6b

Observation 8ff0e854-320f-4a49-b9da-92e317c5fbe5 · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.900333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.900333Z digest=sha256:c36a45f4b517a5d6dfa8e33dd17427d1362ff56a60d332dc4fa80eb588e9280f

Observation f2e8ef3f-266a-4633-810c-aa30d6e6b477 · outbound

This paper cites Rico: A mobile app dataset for building data- driven design applications.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Rico: A mobile app dataset for building data- driven design applications

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.586082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.905456Z digest=sha256:8595a3ad3c941cc4428be590810c78f3227e41d24c82f5844e5f59fbc1257443

Observation 16f9c842-52c8-460f-980c-c3003bb7ad60 · outbound

This paper cites Mind2web: Towards a generalist agent for the web.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Mind2web: Towards a generalist agent for the web

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.571928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.910089Z digest=sha256:1a212f997ade0ae2fbcc87ae406d5196dc0bb049b0348b387c5ea0f5ed5ea84b

Observation 9161f461-1e60-4f8f-ad7d-c8fab868559b · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.919795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.919795Z digest=sha256:b8ac2202efc93b1a417349c441527b09c510bc682f7cbd1e3b03f0a433ffa1a1

Observation 1ecf13a7-a279-4d13-9cbc-f554050ecf88 · outbound

This paper cites Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.558706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.924452Z digest=sha256:f4c4bf942a05eb7c7676b30c997d3d24698e6a109c4c37a5a5cd2618884148eb

Observation 773d8398-3400-446b-b473-f82fb8d2d359 · outbound

This paper cites Enhancing video-language representations with structural spatio-temporal alignment.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Enhancing video-language representations with structural spatio-temporal alignment

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.544854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.929413Z digest=sha256:58e0a100831406a9520577246ffa0a7865db0d3954d12c95035338e86a013345

Observation fb87489d-35a0-4b1b-8583-082cafbd7f0b · outbound

This paper cites Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.934341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.934341Z digest=sha256:3194cd43e4e1fd11e94782ea1854d99dbb19f352aa6e613cd5ad6e77454418f1

Observation 84e459e8-4348-4432-aee6-c3c6fbecdebe · outbound

This paper cites De-fine: De composing and re fin ing visual programs with auto-feedback.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining De-fine: De composing and re fin ing visual programs with auto-feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.529117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.938874Z digest=sha256:ad1eac45efbbf6d56c04bc24ee4c9d6faca3890692febb5785346c02a7b94ff4

Observation 6c3b1587-661a-4697-bd1d-429b24c7400d · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.942871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.942871Z digest=sha256:6535b517bb3c42d38f6d71d8575eb6d28d970ccf6edbee4597d8fb0337132d03

Observation 22c75973-7a83-4169-89d2-a9460907a033 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining CogVLM2: Visual Language Models for Image and Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.947119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.947119Z digest=sha256:3399afa1b111caec84875dfded0dda4445864e979f43d4a8741d1eac203d2a73

Observation a3c0777f-3afa-4d7d-b4da-da34b226b0d3 · outbound

This paper cites Cogagent: A visual language model for gui agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Cogagent: A visual language model for gui agents

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.515180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.951617Z digest=sha256:0335b8662bca588660df69af537786c2d1236cfecc1006da7539312653249230

Observation 1aa0533d-9155-4c66-a343-9d715d165d60 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.955772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.955772Z digest=sha256:fb829cf0c0ef539ea56b9db4c9e7a7af54870e435c18be400cf6cfd667f73ff4

Observation d4cab38e-5f1c-4b08-9d62-1bf1787e5236 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Adam: A Method for Stochastic Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.960093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.960093Z digest=sha256:598b247ad988a00f1942ffdabd54363ac10caafd355eca71b6bfd0bfe39bd179

Observation f944425b-8ed8-4f99-bccc-f632f40237ad · outbound

This paper cites Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.964021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.964021Z digest=sha256:913e06492b320e487595f6cb1df1fa46660786f07a32d34cc533f3b28c4a297c

Observation 129b2dc6-8eab-42b7-a7c6-cc38c7c21a3c · outbound

This paper cites Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.968225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.968225Z digest=sha256:e3c423e8a3d25d01d4b728996a23c4dc8f5399d46737d0f3ccad8c3231dde013

Observation 289ae9cd-18e5-408e-a7a0-98ad1fee55ee · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.973202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.973202Z digest=sha256:b613f279223dc5d0d91c3eb4dfeb45978e2bcbc76e8e6afa9d303856db6bf134

Observation 15d50b90-e9ce-4b7e-a6f2-044a0d703951 · outbound

This paper cites Visual instruction tuning.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Visual instruction tuning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.501267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.977592Z digest=sha256:84682ecf246998b402a96206c7e35c4f13b3fcba9fa9dd21c5f75d4015c8c71f

Observation c24df63a-9a42-40c3-b1d7-218e41340dab · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.981478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.981478Z digest=sha256:11c7efc9446fa04578690dc311bb8721a6d8716ab204ecb0b0dfc32f21b04f47

Observation 40a18b4c-7574-4144-be21-f9f3feecea53 · outbound

This paper cites GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.985979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.985979Z digest=sha256:663cc44fe31f27b3cb9df2b1311de1fffa8e9f942c3cdffcf37a6dd68633996b

Observation 380ec206-46c8-4459-818a-ea52b28e16a6 · outbound

This paper cites Androidinthewild: A large- scale dataset for android device control.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Androidinthewild: A large- scale dataset for android device control

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.487904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:28.990556Z digest=sha256:b3073a53547b057ab6b0ab28f423c4595137c77794cab0e44a7da331296bc0af

Observation 267af3bb-898f-4b25-a1a4-f66b759c0e5b · outbound

This paper cites META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.994614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.994614Z digest=sha256:e3980b64334ca1cbe42ccf07f878b2d93c68f8b0be6db259950d31582d86d2f4

Observation ecc04415-d9a4-4a04-b8e2-a5d56b3acf59 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Gemini: A Family of Highly Capable Multimodal Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:28.999032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:28.999032Z digest=sha256:8ed12d908f5d1da1a2526debd59274fc15c538aabf347286585a8caad4f14fc1

Observation a81b2c2d-c420-4e5b-b61a-f9a6127e8a3b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.003225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.003225Z digest=sha256:18466656c08f3f97e785cd15867b8dfec9f0205927ac8433e5abbad9a90fee92

Observation 6adb5cd8-dc24-41e8-aa01-80ca931d4788 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.007597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.007597Z digest=sha256:49b4fed8aa575421996192961ab411f6c4cb5341cf543b8f741c59f9ca9026c8

Observation d45563af-ff1d-4516-9a54-3295ee6150cb · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.011807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.011807Z digest=sha256:a1330451ff2de211ca00be7c27e820919c16662709c32f4521b614d22f9a40cd

Observation 9a5de42d-6139-49b2-8e43-8918c866aafa · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.016231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.016231Z digest=sha256:cbdaf30e07efa29d6cd01150b9a8ffb40b472a01d5b97d27921df3b19c0841f3

Observation ebddd9d9-b86a-4cc9-ac30-ded666f716c7 · outbound

This paper cites Webui: A dataset for en- hancing visual ui understanding with web semantics.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Webui: A dataset for en- hancing visual ui understanding with web semantics

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:10:29.474038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T16:10:29.021156Z digest=sha256:0dd1c2042ae2136727267986b5972054b0307376e1a2689b4af6edbf7d0d7f1f

Observation b2baf02b-11ac-4e2e-a2ef-d21b6122ef8e · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.025490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.025490Z digest=sha256:76d8df083e098df9ccccd1fc75cf4398aa354aec807b25942605cc4b3f340aae

Observation a0472031-8831-4bed-a82c-bec1764fb727 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.030128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.030128Z digest=sha256:11e052271429866c59a34b1d87e06bca45b5868aa63ef8841891ea5ece14023a

Observation c5859b87-0094-4be4-b9a5-d5c06f4380e6 · outbound

This paper cites Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.034981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.034981Z digest=sha256:da3b895941a990f54a63c7982bb66124ad8003e86ae61c01995563c072071986

Observation 80058f2c-f2a8-43f9-8ec5-975a8065a7b3 · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining UFO: A UI-Focused Agent for Windows OS Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.039295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.039295Z digest=sha256:05fa2d4c221af09ebb18cd41eb79fde5fc9911f6a91baee2186d505ded9dfbce

Observation 94f898c5-0801-4a50-83b5-d594cfead54d · outbound

This paper cites You Only Look at Screens: Multimodal Chain-of-Action Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining You Only Look at Screens: Multimodal Chain-of-Action Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.043380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.043380Z digest=sha256:717f113bc5d9d39e32bdaec2003d3928142c8c5819f6e996d7c2105add548ad0

Observation 7453521a-0ec9-4fb6-aeb3-7a19c1943f9f · outbound

This paper cites AgentStudio: A Toolkit for Building General Virtual Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining AgentStudio: A Toolkit for Building General Virtual Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.048625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.048625Z digest=sha256:6154203157ecedf14fae8abbc2ac6110d854be4f86bd5e7051512bb2f62633d1

Observation ff225322-4c23-4035-a16e-c71e1281f2f0 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:29.052921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:29.052921Z digest=sha256:eb7ed3e82ccc7aac5d927a7f052e535b03dec20f6f4fe6bca70ebff857c5d11b

Pith citing papers

Observation f7565c5c-5909-4277-bacb-4540ca8286ba · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 214

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.769205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:b3d204cc04bed686c767b3466776dc41d960c5cccd754be71d63b1e35e81428c

Observation 47c6f886-57fb-4877-b41b-63b287dd429e · inbound

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities cites this paper.

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:29.178376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:02:29.178376Z digest=sha256:acc0857a8fcf1b97d0407ee695d4d63bae74705607357dab3b40bc5f6618fdbf

Observation 77fc636a-8af5-4362-9d75-3febd07cdabc · inbound

Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills cites this paper.

Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:36.675559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:36.675559Z digest=sha256:51506316cd718dcfeaa9b0e4db3d4b6f86541e34939d771ba6397273dbd05fd5

Observation 3befbff8-0d15-4f3a-8a88-950c2955dd53 · inbound

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? cites this paper.

OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:05:09.289637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:05:09.289637Z digest=sha256:538335320b476a917cac088f2a29b9c6b9ba428eb2c5f8af29c974101bdacb98

Observation 74b247d5-9142-405e-9f6f-6930299855d2 · inbound

Cybernaut: Towards Reliable Web Automation cites this paper.

Cybernaut: Towards Reliable Web Automation Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:21.052024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:21.052024Z digest=sha256:f64d310d7bcb97119e3419ca7917de8baf96e38cc7f92d082ef11f2b00b3322a

Observation 5c57e7fd-9a60-4885-8155-74ba7bfae981 · inbound

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs cites this paper.

How Auxiliary Reasoning Unleashes GUI Grounding in VLMs Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T16:51:32.578220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:51:32.578220Z digest=sha256:9561ff37f3f58f9e9f81657ddfd727ee8a626f543605a6c72aefad5ddf7b854b