Pith. sign in

Paper Citation Record · LEDGER

CogAgent: A Visual Language Model for GUI Agents

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 56 inbound Pith citation observations for arXiv:2312.08914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.08914 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 56 of 56 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:48.765839Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3214e3ab-a280-41cf-8dbd-7e197a925c31 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.446292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:5d6ae839ad1549f68c563c569056306589f08a3e2da074ac67e5d2e8c97f2d36

Observation e6675e01-5d8a-410e-83a4-978dd5e5611e · inbound

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents cites this paper.

SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents CogAgent: A Visual Language Model for GUI Agents

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:09:46.540356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T10:09:46.447508Z digest=sha256:467b00a2a3478c48e98a46c69184fea7e51f73a6bd3acf67587570d41d1e8b05

Observation 639bd103-164a-48a4-bd0d-82cbf0daa853 · inbound

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments cites this paper.

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments CogAgent: A Visual Language Model for GUI Agents

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:19:32.465423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T01:19:32.406859Z digest=sha256:2f9fef861e6b8dc820008c559625395d7f476ddba176b81659e0f9c380500586

Observation 61661f24-5128-4028-9f4a-cab87d879e8f · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites CogAgent: A Visual Language Model for GUI Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:58.990645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:525b4059d7f2e63a5d0d180f0e9a02e349eb135f49911330ddc21f78a94378bb

Observation 033a0815-e370-4512-8311-f765331f23a1 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning CogAgent: A Visual Language Model for GUI Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.935228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:50bba305282c529fd0cf08b2c3c113ae3d5cc5eabdfde2bbd75449f80f107a66

Observation daff75fb-8129-47db-9e38-ec8905cb7402 · inbound

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents cites this paper.

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents CogAgent: A Visual Language Model for GUI Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:06:13.777835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T12:06:13.697487Z digest=sha256:e9e9f3e56f275ffb74969b3b5ca27f30ef983327911d73369191235ac515adcd

Observation d20010ca-3451-405e-8a1e-55c27a43e0da · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.172445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:d3f11df163f8b900b73126783086451d62cb12af297c23009d4305aa121c8efd

Observation 18d3795f-aba7-4893-9d39-29fece85d28a · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output CogAgent: A Visual Language Model for GUI Agents

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.769978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:f6f41f4d153e047938b3f31d0aeb8d7edb96310d2966c7a3595a248eabb9d4e5

Observation e4f5fead-4c87-4d63-bbed-2c9ec1fb625f · inbound

Large Language Model-Brained GUI Agents: A Survey cites this paper.

Large Language Model-Brained GUI Agents: A Survey CogAgent: A Visual Language Model for GUI Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:08:27.968638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T11:08:27.472508Z digest=sha256:cd700e0690385e59d6b8f1d7b95ed9ceac1d503a77beae1496438f7dc62545f4

Observation c6ca42f1-b380-451c-93b0-d5c31f5acc2d · inbound

WEPO: Web Element Preference Optimization for LLM-based Web Navigation cites this paper.

WEPO: Web Element Preference Optimization for LLM-based Web Navigation CogAgent: A Visual Language Model for GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:48.765839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:48.765839Z digest=sha256:59f4738625bbbc8a23d546079a15da00a04887f76f73d668f4dcc6c8775891c1

Observation 1c4bb37b-564a-411c-aaef-10c9a3e84d39 · inbound

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning cites this paper.

Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning CogAgent: A Visual Language Model for GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:38:09.231300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:38:09.231300Z digest=sha256:a8c18813384d043e86d40355f92a6edbe59ce46104c819bc2576387c686e0c65

Observation c0c9453c-ec3b-4a62-b089-95b74fb9fcb8 · inbound

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval cites this paper.

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval CogAgent: A Visual Language Model for GUI Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:21:36.529932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:21:36.529932Z digest=sha256:00ea9b72d46961e8550cd738ff5835e0a0411222ba783988bb1a8f7c0daad8ae

Observation 364ed6cb-90c8-45e8-92d2-da3cb390859e · inbound

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models cites this paper.

Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:58:43.080825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:58:43.080825Z digest=sha256:3ec7f4b1c17002ab9422df1b34b04e5edffb965eb380da4a24713c54a9f80651

Observation 6c829448-aa43-4727-aa99-45ef772a8d61 · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer CogAgent: A Visual Language Model for GUI Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.555350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.555350Z digest=sha256:8d5aea6b03286ab9d48d239576e966fbf169239c428c69ebb3bc0cd69b99e7c9

Observation 217d370a-9e1e-4a49-8939-8182104f7c3c · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends CogAgent: A Visual Language Model for GUI Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:27.630693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:27.630693Z digest=sha256:844b8e5ea5de56093c6d75633fd5eb491ac3af9110a0ce94f0cc16651a8bf241

Observation 03960af4-c56f-416d-8e62-9fa561179df5 · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding CogAgent: A Visual Language Model for GUI Agents

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.437102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.437102Z digest=sha256:e2fe4f5b9dc5bfe592d1de4cc98343f02fa69a994887499a2631690ff0afffc4

Observation cb0d7e34-985b-4a5d-883b-f4a246d12ae1 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:33.815052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:33.815052Z digest=sha256:5ef290f82cee525f582038580f2b53c2179ca00595e4756c1d0fe7465745c32c

Observation c4a373ce-8807-43ad-8fc8-84f984a0d932 · inbound

InSTA: Towards Internet-Scale Training For Agents cites this paper.

InSTA: Towards Internet-Scale Training For Agents CogAgent: A Visual Language Model for GUI Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:24:50.335297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:24:50.335297Z digest=sha256:19437ef618701548058fc3b4e9262df23ddd5be1de49e8190252d4eb103dfa13

Observation ecb5431c-b222-4c9f-85f8-60a5fe5f06a1 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.669886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.669886Z digest=sha256:ec37ab7f38d0ab3d172c92e1b44478011bd1b94cecb238e32897a6e479ba6a3e

Observation 485324ff-eeda-4a5d-8817-79a990391ed7 · inbound

TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents cites this paper.

TRISHUL: Towards Region Identification and Screen Hierarchy Understanding for Large VLM based GUI Agents CogAgent: A Visual Language Model for GUI Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T06:01:31.626071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T06:01:31.626071Z digest=sha256:93d19724a92ae6d362fa83525f8ae0ceb929ab23dfb6fa4aa6c99a33402312e1

Observation ae727bc1-7c64-44bf-b741-9d442477c5da · inbound

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents cites this paper.

Digi-Q: Learning Q-Value Functions for Training Device-Control Agents CogAgent: A Visual Language Model for GUI Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T20:57:48.425527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:57:48.425527Z digest=sha256:2507b7326bcac022e1a2f3cc8837b22427941e0541a02ac88231bd3d7527709a

Observation 53784e52-3283-4c3c-ae0f-0bb1fb454dde · inbound

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation cites this paper.

Aggregated Structural Representation with Large Language Models for Human-Centric Layout Generation CogAgent: A Visual Language Model for GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:04.556039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:18:04.556039Z digest=sha256:966bec75fba4ba3eceb90411aa9f332c112220f39d92dd2cb1d9b6aba13dc436

Observation ffb37659-bcbc-4cef-8a15-a406b0667597 · inbound

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models cites this paper.

DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:43.994161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:43.994161Z digest=sha256:16f89fe0aacee37251eb66379e56d834b7321d0804769f19c2d351d56c1910f6

Observation c4fb6c79-0874-4ebd-817a-ace03b950ed7 · inbound

Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation cites this paper.

Exploring the Potential of Metacognitive Support Agents for Human-AI Co-Creation CogAgent: A Visual Language Model for GUI Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:13.264774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:13.264774Z digest=sha256:00bbc229518876acd81ff6a96bc148e815c5da478f4145e9f1ea284d9d827d7e

Observation cc4f8ab6-75d4-4a7d-bf8b-aa198f685dbe · inbound

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding cites this paper.

ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:44:51.212777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:44:51.212777Z digest=sha256:ec521c2ac8ccdb8f6f5bc9422ea08ba8b74f4a3d290aa75ba5afbd283362328f

Observation 500ea2d7-0380-4833-8837-d97ded882cd3 · inbound

Mobile GUI Agents under Real-world Threats: Are We There Yet? cites this paper.

Mobile GUI Agents under Real-world Threats: Are We There Yet? CogAgent: A Visual Language Model for GUI Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:02:07.944025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T06:57:25.257775Z digest=sha256:4f06b82954ac143a041d353583da0d56cf34f25f101c7fe02b3c3b34f4878d60

Observation d0f6657d-6789-4ca2-a6f7-4dcdb6d7b6c8 · inbound

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding cites this paper.

GUI-G$^2$: Gaussian Reward Modeling for GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:11.323765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:29:11.323765Z digest=sha256:4eb45cf646b3a005378d86b23b9413b6ab78c54c579063475fffed3264710ded

Observation c8769967-487a-4cb0-973a-086bfb325e7f · inbound

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation cites this paper.

Screen2AX: Vision-Based Approach for Automatic macOS Accessibility Generation CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:08:35.405743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:08:35.405743Z digest=sha256:7b4df2a43ffb3e17eabf3517e94d20506acda97666ff222c878fcbf5c9713e9a

Observation b9019cd6-d586-44bf-9fd8-a4568a2e5b9c · inbound

Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement cites this paper.

Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T01:01:07.165776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:01:07.165776Z digest=sha256:d5060f0ccd7d8d6dea27355f8c79c47c5b7ceade80fa45cd93a8c55d94962166

Observation 174dd392-0872-481e-9294-fc014b238dca · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience CogAgent: A Visual Language Model for GUI Agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:48.651764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:48.651764Z digest=sha256:64b0e532b1a1572b10376e90e5f79b7073f922c554b31c0890d4dc1b4d3884b1

Observation 0aeabdef-01b3-41ce-91ab-3bbc1577c9e6 · inbound

Cybernaut: Towards Reliable Web Automation cites this paper.

Cybernaut: Towards Reliable Web Automation CogAgent: A Visual Language Model for GUI Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:20.465838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:20.465838Z digest=sha256:29c711858db491a7e9ec98faa175e4a889c7488d929b606b1e9fb7e50ee31d93

Observation 0f880a2d-6315-4c62-ae82-8b20b54c5721 · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.628996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.628996Z digest=sha256:8d7b2a9a419fc895ae89e02fe302041a3947b517bb37b178bd38ec82ad783dca

Observation 56f30723-c351-4607-9e16-bbe77c31f994 · inbound

MobiAgent: A Systematic Framework for Customizable Mobile Agents cites this paper.

MobiAgent: A Systematic Framework for Customizable Mobile Agents CogAgent: A Visual Language Model for GUI Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T13:34:59.368315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:34:59.368315Z digest=sha256:94d4028664d689f0979e2b955f85677f20fff71b985305debd6c4eaf2159ce83

Observation 54e9756b-4558-43e2-8a15-858c219cc315 · inbound

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action cites this paper.

UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action CogAgent: A Visual Language Model for GUI Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:03:38.142743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:03:38.142743Z digest=sha256:654cee96cd4454aacf9bb238b7fb59ce4997222d8b53ae6722d814005dc7a948

Observation 3be052ac-8eed-4fad-9a52-bf977b2ebc20 · inbound

Grounding Computer Use Agents on Human Demonstrations cites this paper.

Grounding Computer Use Agents on Human Demonstrations CogAgent: A Visual Language Model for GUI Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T23:06:04.882108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T23:06:04.882108Z digest=sha256:7894ceefadef5e4bd91f60380054105477bbe8cabc8bfd4318c6362d83a4df87

Observation 0ed4a0e8-0cda-4841-95b3-4d180ee4915f · inbound

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL cites this paper.

GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL CogAgent: A Visual Language Model for GUI Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T20:51:41.781289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T20:51:41.781289Z digest=sha256:28969c0ecb5e656a22eb0de3dd82a257b923fab969cbf74c0a65f067c952123e

Observation db48a1f3-5977-4056-9f33-511c1ac79750 · inbound

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting cites this paper.

Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T15:29:31.834566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:29:31.834566Z digest=sha256:170d25a699387eae61acebf2928923b283cb45e7ad714e95d63413418e701e62

Observation dd3a4c5c-dccb-4200-8e20-0c5f3b88b3e8 · inbound

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web cites this paper.

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web CogAgent: A Visual Language Model for GUI Agents

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.721350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T18:00:34.401698Z digest=sha256:30c9929dd7954a332e7adb69277330d05d9c98ff5c1ed3e714e1242d62fb2467

Observation b3b6f78c-33cd-47bf-890b-39f1bd64ead3 · inbound

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding cites this paper.

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding CogAgent: A Visual Language Model for GUI Agents

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:10:28.874668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T14:06:55.472857Z digest=sha256:819c060947f18781fc3e3f6176b3633fb4e7a7cf361bb67d3c9067d1721824bc

Observation 8a73181f-a9c7-45ae-99e5-6d60903a08d1 · inbound

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory cites this paper.

OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory CogAgent: A Visual Language Model for GUI Agents

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:25.492638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-07T10:45:48.976501Z digest=sha256:7c900d70025e6a9fb07a01ff065f35cdf58891d92387f52525b784a3ebd7ded0

Observation 85007b8e-8431-49fe-82af-c599ef7c4e98 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:07:51.413266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:05:36.511150Z digest=sha256:eade0e998c431b67b2132e48c6a04f6a402a6a3329ca9c9a1914ef0813bed23a

Observation 27da449a-793a-4628-95e2-045219275cd9 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:59:48.175276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T05:59:44.669877Z digest=sha256:5d1bf9f03b027468200f4694fbd719c21e7602da817c833147fb4f12c8b0a480

Observation 94636b68-cf0c-4e6c-ae31-eed6da1a0af7 · inbound

MMSkills: Towards Multimodal Skills for General Visual Agents cites this paper.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.557252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:493a585bcee7448a406f0c4148a84faa9749df912c42b946084a71c3237780a0

Observation 7c800237-0192-4771-8bc4-b9d0d66db98d · inbound

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis cites this paper.

ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis CogAgent: A Visual Language Model for GUI Agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:04:37.662147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T10:57:59.013811Z digest=sha256:6846f837e99bda6a13a0b171b1663038a33d1e5b958b4da3f78c0995878244b5

Observation d44059e2-596b-4f0b-b5e0-6fb567de06a8 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap CogAgent: A Visual Language Model for GUI Agents

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:02.060945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:7ccb3b43e9e8d9a04696ae289525ab3af6981e3f12aae5a0bc6f80189872934f

Observation 546643e6-e15f-424f-94c6-5f61d0749a71 · inbound

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark cites this paper.

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark CogAgent: A Visual Language Model for GUI Agents

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:13.153036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T07:38:35.829605Z digest=sha256:5c321681530724447f6e1efaa326f9cdf2f411972ad69a19efff92cd16b00848

Observation 65db81a9-6606-4ac4-a108-bae0d9e901bf · inbound

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning cites this paper.

GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning CogAgent: A Visual Language Model for GUI Agents

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:44.696258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T22:52:04.755523Z digest=sha256:ba86300a33eae29556dd0cad851f754b21e75489ed921cec0cf3ade8b48cb33d

Observation 749f98ca-d7b8-49b0-9933-53d2beb899bc · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation CogAgent: A Visual Language Model for GUI Agents

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.013381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:64c19fda81121f211fca98067539d8200c7692fe9a030b97e9ae8b539f91869c

Observation 84ef5e59-b519-46dd-bf4d-9381a206014c · inbound

ProCUA-SFT Technical Report cites this paper.

ProCUA-SFT Technical Report CogAgent: A Visual Language Model for GUI Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.718222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T03:18:04.149281Z digest=sha256:6102ca917d5cfb742a6683a71489bc1bb0309abba1377daf24f56761baed984f

Observation bd8d87a1-6bcc-4cfd-84ef-01068cbd907e · inbound

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents cites this paper.

MIRAGE: Stealthy Visual Prompt Injection for Vulnerability Detection in Web Agents CogAgent: A Visual Language Model for GUI Agents

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:08:58.345016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T00:54:00.944706Z digest=sha256:f6b8d81e5852b8b4a0d9d39cde0bc970ce92650e4f41bb888b307c02cfc8203f

Observation fe492bcf-aa7d-4260-9aad-21ab9317c5e1 · inbound

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction cites this paper.

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction CogAgent: A Visual Language Model for GUI Agents

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T07:54:22.413201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:48:01.719339Z digest=sha256:313cbd4f939e1108da882bd322c0029004a761ef950ef10dff5b47bca86fff4e

Observation 959a6d74-e4cc-420b-8379-73eb220a464b · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.193108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:1880a09e376c8e0f257085daec275dfbff029fee2651c56c9d9a1a2be2152f3e

Observation a12d611d-2687-48b6-b2bc-a2aa74e42147 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks CogAgent: A Visual Language Model for GUI Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:5e34c3b45c02aeb6c2d2f0303563be0dfa986e20fef0f7ff166d64252e56b6a6

Observation ab5b6383-dd99-419b-b36e-d40de6cb68c1 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report CogAgent: A Visual Language Model for GUI Agents

Reference 110

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:e999592bb167c1382b61a14f0904abc686bf984bbf475cbc1a2b01bd69922df0

Observation 3e873257-3ded-4458-b51d-c2dc656c60c2 · inbound

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models cites this paper.

Capturing Token Tendencies for Training-Free Token Pruning in Multimodal Large Language Models CogAgent: A Visual Language Model for GUI Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T10:42:21.217308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:42:21.217308Z digest=sha256:8dde884a43049ba49dea4f59c74c9ad3876db8bdc934500520e9ffba9ac3be3d

Observation 19c92063-4d6b-49e1-96ae-3daa64838819 · inbound

AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents cites this paper.

AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents CogAgent: A Visual Language Model for GUI Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T21:40:39.900263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:40:39.900263Z digest=sha256:678d2062a0138db390bb898ca58dacc42015f10e7d9116c18d4cde81d99c65e4