Pith. sign in

Paper Citation Record · LEDGER

MMSkills: Towards Multimodal Skills for General Visual Agents

As of 14 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2605.13527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.13527 v3

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:40:24.948959Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T21:28:58.598336Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact13
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch24

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27f60dd1-2323-4e7e-8558-d4e2245774ab · outbound

This paper cites Agent S: An Open Agentic Framework that Uses Computers Like a Human.

MMSkills: Towards Multimodal Skills for General Visual Agents Agent S: An Open Agentic Framework that Uses Computers Like a Human

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.539803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:28978f75141b3e036a5c800370d085f3d3b1c1446748e661dda27feed5145b15

Observation 3404e158-0908-4681-bd78-52d75c500e8d · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

MMSkills: Towards Multimodal Skills for General Visual Agents Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.553070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:ae5befa80e2a38dbb873d6b477dc699a397b0e76185923a59d4d86d5ada0446d

Observation 7c5ea0d4-d852-4801-beab-83f5d2c2ce0d · outbound

This paper cites EvoSkill: Automated Skill Discovery for Multi-Agent Systems.

MMSkills: Towards Multimodal Skills for General Visual Agents EvoSkill: Automated Skill Discovery for Multi-Agent Systems

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.564993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:aad0e8606ce4458f6a491fbe619802723296d4951e2c04ba8704ac39347023cd

Observation bcc50edf-73f8-4ac3-9788-e5fd9fc07dc4 · outbound

This paper cites Qwen3-VL Technical Report.

MMSkills: Towards Multimodal Skills for General Visual Agents Qwen3-VL Technical Report

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.531773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:05a2b9f369b26d510a34bdebb996944e70343935ba593cda404af9f336767856

Observation c9ba6989-a8b5-4648-9d8e-c9d592447fdb · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

MMSkills: Towards Multimodal Skills for General Visual Agents LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.567653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:47f32074be746efb3bb0cd2734700867c1ca5356067b3f4c790c7cdececa655c

Observation 3392dde1-a5a7-4dff-9ad1-10d98d94482b · outbound

This paper cites Cua-skill: Develop skills for computer using agent.

MMSkills: Towards Multimodal Skills for General Visual Agents Cua-skill: Develop skills for computer using agent

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.570747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:4a36082f4bd84c8ad0be50bd2f2e5dee1a9d50fa5a59063861ac9718d71222c6

Observation 1ac47709-125b-4b0b-a94a-6c1fe1b96c22 · outbound

This paper cites SeeClick: Harnessing GUI grounding for advanced visual GUI agents.

MMSkills: Towards Multimodal Skills for General Visual Agents SeeClick: Harnessing GUI grounding for advanced visual GUI agents

Reference 7

Resolution
metadata mismatch
doi, observed 2026-06-30T21:35:04.028125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:9b986ea9bed34de637be1225b9d8829d2fb2d1d72ce2111d7f17c29a894bacc3

Observation f9c180a3-5e0f-49ba-899a-d1423645518c · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

MMSkills: Towards Multimodal Skills for General Visual Agents Mind2Web: Towards a Generalist Agent for the Web

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.573142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:0a801b9498e1fcf95c4060359f88aade84a97f0e193d5455224c92f3c0f47282

Observation a5d0bb76-3bb8-4243-98c6-88069d3677a3 · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.554540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:71cfcc5924205cd3e418c9399ca8cdb4cb60a209ada5e44384d76b7538ff5bd1

Observation 2d08272c-fa46-4bac-983d-5a6ae8850db4 · outbound

This paper cites Webvoyager: Building an end-to-end web agent with large multimodal models.

MMSkills: Towards Multimodal Skills for General Visual Agents Webvoyager: Building an end-to-end web agent with large multimodal models

Reference 10

Resolution
metadata mismatch
doi, observed 2026-06-30T21:35:04.018795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:8c6f8c15a4a6cb65d14858ba5de1a8e1476e4eea52ae3ced1885847e5268a958

Observation 94636b68-cf0c-4e6c-ae31-eed6da1a0af7 · outbound

This paper cites CogAgent: A Visual Language Model for GUI Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents CogAgent: A Visual Language Model for GUI Agents

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.557252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:559e555b2bf3d218bd6ecfd4570bce9431982bd94ca5b171d3f5525b7aca642c

Observation 1d368f40-3266-41ea-8cdd-8b9fb6fa95df · outbound

This paper cites lmgame-Bench: How Good are LLMs at Playing Games?.

MMSkills: Towards Multimodal Skills for General Visual Agents lmgame-Bench: How Good are LLMs at Playing Games?

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.559898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:13677b085a5e8c586494e51ac5880ae190433f8b84eb8b4a63cbb0d91875f2d0

Observation 0287352e-4c26-41e5-ad64-53153b595a65 · outbound

This paper cites URL https: //doi.org/10.1162/NECO_a_00393.

MMSkills: Towards Multimodal Skills for General Visual Agents URL https: //doi.org/10.1162/NECO_a_00393

Reference 13

Resolution
verified exact
doi, observed 2026-06-30T21:35:04.020654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:d462be001846b6776f2417af6d4cb91e04fc5183ecf2be0255a271bd93908107

Observation 979ab9c9-eef2-4982-920c-ee23418528e0 · outbound

This paper cites XSkill: Continual Learning from Experience and Skills in Multimodal Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:17:26.572211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:48e6ad7100bdcd34351e66c1a7f217ac9b0987e77eaff80b5791ec06ff38ca9f

Observation 4087e90b-091f-4724-93f4-c624695fe047 · outbound

This paper cites VisualWebArena: Evaluating multimodal agents on realistic visual web tasks.

MMSkills: Towards Multimodal Skills for General Visual Agents VisualWebArena: Evaluating multimodal agents on realistic visual web tasks

Reference 15

Resolution
metadata mismatch
doi, observed 2026-06-30T21:35:04.034096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:64f7d9985487caef739a834048e0280806db34c24c56ec8ec0de86543314df35

Observation 01b9da5e-81cb-4be0-abad-2355726a4b21 · outbound

This paper cites ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions.

MMSkills: Towards Multimodal Skills for General Visual Agents ImmFusion: Robust mmWave-RGB Fusion for 3D Human Body Reconstruction in All Weather Conditions

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.025940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:d08781b02a84420995f61aee51dc1c1080f8e49b9e990b759b85c3b805eef68d

Observation 542aed55-042d-4ba5-a172-7ad1f4a18820 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

MMSkills: Towards Multimodal Skills for General Visual Agents Lost in the Middle: How Language Models Use Long Contexts

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.023262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:283770deeece63e032e2bbe5dce760d09c90c77a189ef65bdabd6fefe90e78b4

Observation 5754cc06-b451-4cb4-a214-f528d2cfef92 · outbound

This paper cites How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings.

MMSkills: Towards Multimodal Skills for General Visual Agents How Well Do Agentic Skills Work in the Wild: Benchmarking LLM Skill Usage in Realistic Settings

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.543425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:b14f5ca3755e0623505cc9a2b5a5c853f91ac42c805c8ff920d9ac72823b6dd2

Observation 1edb1189-7738-40a4-9a4b-387635fc2c22 · outbound

This paper cites OmniParser for Pure Vision Based GUI Agent.

MMSkills: Towards Multimodal Skills for General Visual Agents OmniParser for Pure Vision Based GUI Agent

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.549327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:4fceba8c562259eef9dcad1529755c8463c7cbd55b4d48c604ce0b0c69f7ddf2

Observation 20ce9619-cb9a-4aa5-a70d-2ab5746ca70d · outbound

This paper cites SkillClaw: Let Skills Evolve Collectively with Agentic Evolver.

MMSkills: Towards Multimodal Skills for General Visual Agents SkillClaw: Let Skills Evolve Collectively with Agentic Evolver

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.580530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:61c9940b42edc07b1a7113837ee293fdb38edb3e01f685a383daf33ef24dacb6

Observation 542e943b-0bd4-46d0-a554-da29ca0c95b1 · outbound

This paper cites URL https://doi.org/ 10.1017/CBO9780511811678.

MMSkills: Towards Multimodal Skills for General Visual Agents URL https://doi.org/ 10.1017/CBO9780511811678

Reference 21

Resolution
verified exact
doi, observed 2026-06-30T21:35:04.030066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:3755d1211f84351a312f0db1c8e515bfce30bfbe61e9cb7d6cefbc5225b111d6

Observation 17e48366-d340-4fb1-a2c8-84c073bb9bbd · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

MMSkills: Towards Multimodal Skills for General Visual Agents MemGPT: Towards LLMs as Operating Systems

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.583018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:67c7470f035841fdc60ebbe92af4d07a8a35ab1541b9635a0901c2e754bd774e

Observation 788c8ecb-23b7-4bee-bc96-179de575995e · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

MMSkills: Towards Multimodal Skills for General Visual Agents Generative Agents: Interactive Simulacra of Human Behavior

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.562519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:1a06e419595f2a2f756e0c93490c8d5c127c14a06baace4218b1daf609fa395e

Observation 8fb550e3-a08d-4eb1-aa7c-44136aaa74aa · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.540664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:ba937d1cf2235c66bfd762cfbbd472dafaa828bfeed9598c713e51a2e0c4976d

Observation 9405978a-5fbe-4e3a-90a0-cfe5cee7c255 · outbound

This paper cites Android in the Wild: A Large-Scale Dataset for Android Device Control.

MMSkills: Towards Multimodal Skills for General Visual Agents Android in the Wild: A Large-Scale Dataset for Android Device Control

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.546582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:a6da26ed04c1bb089829875d4929c35601ac71a77619e58c8c31f525b3df786b

Observation b8502d64-b27a-4b45-b0e2-b68c71a1a255 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.578198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:a218590dfb45c5db507eacea912f83038730b5920cd2369c61d14f48c29aba85

Observation a372cebe-cfda-4c19-a95a-4794b8b719e8 · outbound

This paper cites nips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html.

MMSkills: Towards Multimodal Skills for General Visual Agents nips.cc/paper_files/paper/2023/hash/1b44b878bb782e6954cd888628510e90-Abstract-Conference.html

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T16:53:57.384410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:7a0a74cb5711e4064a7ebde57ecbae55cc3a3ba48cdbe73438ae7aceaedf54a5

Observation 47b3dbea-0a21-4aac-a219-8fdd899775fb · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

MMSkills: Towards Multimodal Skills for General Visual Agents Kimi K2.5: Visual Agentic Intelligence

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.032447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:db28c5840bfa9913e00d2766361f83eb2b0fad3373369c71b446d015b0c6e59f

Observation d6c43618-8764-4b86-98a2-895dde225359 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.531961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:05f62c4739e18a8023ab6e8db88e61e6b728ef0caa11111f1a9d91ea1a45a723

Observation b45f1b66-cccf-415e-8da9-d36a804694f2 · outbound

This paper cites SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning.

MMSkills: Towards Multimodal Skills for General Visual Agents SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.575834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:7dacbe858a837f951647c6c1aa5d17c2c4d1ded7d1f005161dc1118d223b3758

Observation c2f5d302-2ccc-409b-8ec0-593dd6d874ea · outbound

This paper cites Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills.

MMSkills: Towards Multimodal Skills for General Visual Agents Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.529397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:eeac5e6666242fd89e74d30ffc9ef50a597cf22efa3e8ce71f7168b5e09ac19b

Observation 65319806-9555-4e7f-86a3-80378cb86217 · outbound

This paper cites Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward.

MMSkills: Towards Multimodal Skills for General Visual Agents Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.534667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:c422a013464bb126776854f44eb925b4559069ab73655558ac6d45e093a279dd

Observation 6eb4e3bf-a14c-45f1-82a2-4d718435a0b5 · outbound

This paper cites DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.519678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:372dd460d08eb52e636dc24074fb58c8c2c114ed941d46a5ec391238d80879d0

Observation 4832a7eb-97d6-44da-bc89-bc085fb57c55 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

MMSkills: Towards Multimodal Skills for General Visual Agents AppAgent: Multimodal Agents as Smartphone Users

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.526584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:e51c479e53232656ccd958e7412fec32fe04bc4200904becd54e60b545e7f3bd

Observation 1c3874fc-d936-40cb-b361-4736247cd03b · outbound

This paper cites DREAM: A Dual Representation Learning Model for Multimodal Recommendation.

MMSkills: Towards Multimodal Skills for General Visual Agents DREAM: A Dual Representation Learning Model for Multimodal Recommendation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.522570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:55e3b156babd519505b4c9eb380a9d495be39426ef52f28883c9ad104d3bc5a1

Observation 00a25186-d3ee-4694-b252-d75c54fde55e · outbound

This paper cites Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su.

MMSkills: Towards Multimodal Skills for General Visual Agents Boyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun, and Yu Su

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.537817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:d47f38ee43202ed44d7007b61a7835e28ce882b13fa9d58423a8d4c5c2b2ee7d

Observation 83b02c05-c62d-40b7-aca2-02fac7f769a8 · outbound

This paper cites GPT-4V(ision) is a Generalist Web Agent, if Grounded.

MMSkills: Towards Multimodal Skills for General Visual Agents GPT-4V(ision) is a Generalist Web Agent, if Grounded

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.507748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:839bee0594e79fc0b471242552475d46f3f3eed6dcdaf6dae6f7639f08dd65dd

Observation 050be9a3-a594-4c99-b9c3-26a3f3550514 · outbound

This paper cites SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills.

MMSkills: Towards Multimodal Skills for General Visual Agents SkillWeaver: Web Agents can Self-Improve by Discovering and Honing Skills

Reference 38

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T21:35:04.537066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:9d8551652b51ece8b0c951ce70ea20040b83e1601c1a24d097e2d38ed93c54af

Observation d60c3f31-2c53-4485-8036-1ed7c73cb0cb · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

MMSkills: Towards Multimodal Skills for General Visual Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 39

Resolution
malformed identifier
local_arxiv, observed 2026-06-30T21:35:04.501725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:94d970ac752521d393b65fc6c4aed6b50b9c544e561d4b176bb94c257c88f75c

Observation c528b9a5-e8c8-4dcc-8d27-20012df7a473 · outbound

This paper cites Recent LLM agents have made skills a practical interface for storing and composing procedural knowledge in language-conditioned environments.

MMSkills: Towards Multimodal Skills for General Visual Agents Recent LLM agents have made skills a practical interface for storing and composing procedural knowledge in language-conditioned environments

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T16:53:57.382533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T21:31:20.079403Z digest=sha256:17aa1e421f20eac4652c600785e981cb29afb61ddc412f8546cad69e625bba5d

Pith citing papers

Observation 702efe0b-ac8e-46d6-8d36-888bc44bc171 · inbound

VISUALSKILL: Multimodal Skills for Computer-Use Agents cites this paper.

VISUALSKILL: Multimodal Skills for Computer-Use Agents MMSkills: Towards Multimodal Skills for General Visual Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:28:58.599988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T00:35:02.162386Z digest=sha256:7661a163cd15bd83ea571e2632d59dd5c69eaa4045a7718283d48ea1084b4da9

Observation d40bb2e2-fc90-4791-a14a-fdce8722497c · inbound

Progressive Agent Skill Generation via Reinforcement Learning cites this paper.

Progressive Agent Skill Generation via Reinforcement Learning MMSkills: Towards Multimodal Skills for General Visual Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:05:27.821099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:05:27.821099Z digest=sha256:aee189af3838421f6f96e6e211ba0c3ff5a34d59f47c986cda4c5f10c2890376

Observation fbc7bbd1-c0e6-4a53-a802-1110597c64e5 · inbound

SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation cites this paper.

SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation MMSkills: Towards Multimodal Skills for General Visual Agents

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T17:40:24.948959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:40:24.948959Z digest=sha256:8fff43d44f61f7926446c6bbea37a4bc2115fe220e55c54b945f68cdfa47c08f