Pith. sign in

Paper Citation Record · LEDGER

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

As of 16 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 4 inbound Pith citation observations for arXiv:2505.24878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24878 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:27.016316Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T14:45:56.506288Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:56:19.580749Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved38
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d9246e3-756a-499c-88e0-38d0731381b7 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.058110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.058110Z digest=sha256:2d0b85ecdb2ea6131e1cd8a419baf7c6bcc139b7aa03e5c469e0277762dc9f1b

Observation f1fde33c-9802-4c66-9479-4a51cca99e45 · outbound

This paper cites Claude 3.7 Sonnet System Card.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Claude 3.7 Sonnet System Card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:30.038832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:22.148677Z digest=sha256:3e4fa5957551e8a44bca4edfad3517206ffe20832d556bc1220ad3639243af54

Observation 8cd90733-1730-4365-a1f1-1da81cf6a2df · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.240512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.240512Z digest=sha256:0cbd778ff8f050cca849506f0eb323efa15df4d00dd904dfa3578477587dfca0

Observation 8f992786-6b47-46f0-b163-09ba2149e385 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.413093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.413093Z digest=sha256:4b4ca29fd857ddf3adea1e3fa2f738afc8c3bc71c0c80d1bea6d39bbfefaf795

Observation a7884ebb-dedb-46c9-b75f-d74ce5e0840c · outbound

This paper cites ReTool: Reinforcement Learning for Strategic Tool Use in LLMs.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReTool: Reinforcement Learning for Strategic Tool Use in LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.513582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.513582Z digest=sha256:a38a3d73cd9425153c05e1b64575b32901cbee7e76a6be18ecc9f8261fa84c88

Observation 92da3802-4361-4744-80a0-7b8dd8df97c1 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.644201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.644201Z digest=sha256:11e7afccddbcdae3ae12f3e5fac607641794d34d19c6d6b16a1df76d09d13d1e

Observation c7f316e4-5392-451e-afc9-0b28dc1118ec · outbound

This paper cites Gemini 2.5 pro: Our most intelligent ai model.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Gemini 2.5 pro: Our most intelligent ai model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.915407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:22.722596Z digest=sha256:9a5ed87cf767273bc8020330196f046fd52ff880d3874f4825ae7a38b7b6240d

Observation cb064953-eadb-486c-bd05-5b0e3ab59606 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:22.878232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:22.878232Z digest=sha256:3749c4e9170bc7e4f1f1c34c0042d8be98c7fe7c18d46051a34400142c2c1129

Observation f17ade13-c1bf-4358-9893-8e99c5fef9a9 · outbound

This paper cites Making the v in VQA matter: Elevating the role of image understanding in visual question answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the v in VQA matter: Elevating the role of image understanding in visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.588992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:22.979257Z digest=sha256:dc664da722f025543bec54f5fe5df9eddc3d029551ddf300ecfbfc4d3c7394d9

Observation 4ac3070e-20df-47bc-9ca2-4d8c227e0e65 · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.076403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.076403Z digest=sha256:450c82696fecbb63a35fa47715e2d08d6177e4f913cdc689e7a8eb46a98304b3

Observation 5223b282-d58a-4cd3-a7d7-12bb6f221b0f · outbound

This paper cites PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.180088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.180088Z digest=sha256:50bab87d469d39e4986669d5b2956531f834246ee653e351a8d9467358a9f52c

Observation 4d5a1ea1-65a1-4371-b8c7-ab92b984136f · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.288396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.288396Z digest=sha256:465eb8e977d2b1d443101bd2f6ba085150ecb660af1f959176d283b75ee4e36d

Observation e3d8a4ee-d7b6-4213-a4f5-a3f1bce89de7 · outbound

This paper cites FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.358620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.358620Z digest=sha256:abf50dac62908885c3bee14d452035123e74acc54d566009780bd326576a82e1

Observation f02b9289-7554-4051-ba3c-2161f2c635a1 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.446173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.446173Z digest=sha256:ed025f91108890b0783da421c4a3525e8abf67e90cf627df5cbd5e2ebaad2253

Observation 7c9e3463-166b-4613-881a-2dbfec41ae86 · outbound

This paper cites VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.507875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.507875Z digest=sha256:9819f01b8177d501ffb442ced68d2c993fb79112462bb89deff0fa4907f20f42

Observation 4d802964-4990-42c3-a3a9-833eac9fc557 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.619007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.619007Z digest=sha256:09670bf1cc05c8a76e552c5b8318186525006985ae2ba6c5aefb75048e4be716

Observation f7e5c7f2-04d0-407f-88fc-59ac5af16301 · outbound

This paper cites ToRL: Scaling Tool-Integrated RL.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToRL: Scaling Tool-Integrated RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.665196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.665196Z digest=sha256:4d9430a1beca6561f32a1b2d86f4a6cf42d491fcc0d3993703c9f0b39405ad49

Observation 6adc218f-cc4e-4935-9596-8e6def33a041 · outbound

This paper cites Visual Instruction Tuning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Visual Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.747374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.747374Z digest=sha256:681698b9831afb9b3661a9bce47993e641467c5d34090bdf68dcfd3bd122d8a8

Observation 085bd903-297a-4553-a88e-6c011123e860 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AgentBench: Evaluating LLMs as Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.821960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.821960Z digest=sha256:59f3afe0d7694cf2bf523a5f3b5d32f4b790d9f7236e8952c81e012c0ea70a54

Observation 742ae582-6b3c-4111-ac23-744c6737e7d6 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:23.897548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:23.897548Z digest=sha256:04df7fc9ac4a842177747b4d35090630c8fa6a88f3d0090cd924ae49c5ffc9d2

Observation a88e7b80-dd04-41b1-a90f-4c9d997186d1 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.009766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.009766Z digest=sha256:17ccc4efaf54c36a932fe4c16db36eb50125d3969b6e5a2d44f67253ca895462

Observation 8cea61a5-de58-4ee5-a8a8-9174f04cd906 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.419922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.090500Z digest=sha256:871da9736bed148d585f4f9a4f3125e1b41db5929220fe62885bb2aeaaef71b1

Observation 53e93438-613d-47f7-bda1-d00dd6bc3ef8 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:16:29.297288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.177550Z digest=sha256:38999a3f4a38fb7508cb47d890f2fe18c8d73ebeb3343393875f6787cf41fd03

Observation bcd00016-d170-4464-abb1-30000aab1cf8 · outbound

This paper cites Browser use: Enable ai to control your browser, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Browser use: Enable ai to control your browser, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.192433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.252099Z digest=sha256:25d3d538c7968a1aa1199d09ac45d0adfb7198380f6414e3c50407a480aa1a92

Observation f3a0d1e3-f580-453e-beb9-1dfc4044fa32 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebGPT: Browser-assisted question-answering with human feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.378965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.378965Z digest=sha256:e4ef0f89b21a1c181754b15929ac39beb702331e21d23449d6393fdda5cdcd3f

Observation c369baf7-f509-4b3b-9b3e-2e13c5623797 · outbound

This paper cites Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Deep-captcha: a deep learning based captcha solver for vulnerability assessment, 2020

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:29.084204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.467672Z digest=sha256:8ba40f7249b3477b8587dca7dee287b2bdfbbb102f920e458d5f20c7bd7247d8

Observation cbb6873c-85fa-4b90-b060-8caa652f0848 · outbound

This paper cites Openai o3 and o4-mini system card.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Openai o3 and o4-mini system card

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.977538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.581003Z digest=sha256:40c738444959a5ea39caacc3644ae4f4272c23f2f8cad37bd1eac2c0d5befe52

Observation 00261942-5dbe-4968-87d8-283ee522ff59 · outbound

This paper cites Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Aleksander M ˛ adry, , et al

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.674104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.772851Z digest=sha256:2b5052b6756931a3fd32efe1f844418b4502515bcba0240bf07319a2ee8ebd8e

Observation 12e7a838-09c1-443a-ac54-c9d83c113092 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:16:28.841033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.679222Z digest=sha256:2badb3da4f15f91a9a67cd15876956dcebaddc4035604baa586eabf8580b3459

Observation 29736ada-6fbe-413c-bfc1-7cd7b3432b3c · outbound

This paper cites Symbols of One-Loop Integrals From Mixed Tate Motives.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Symbols of One-Loop Integrals From Mixed Tate Motives

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:24.914755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:24.914755Z digest=sha256:89c863fc98ec2b5e468ecf9f409a0384dfaa24999408a0955d4f19c8a79b8b31

Observation 01e7ca81-d3a2-4dc8-af91-f8ea9b4dd923 · outbound

This paper cites Autoplan: Automatic planning of interactive decision-making tasks with large language models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autoplan: Automatic planning of interactive decision-making tasks with large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.504648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:24.857751Z digest=sha256:6ac6f78daf320d21774be292f66f9456e20c8bcf32c856b8d6af3e3e5d55c3b9

Observation 602fb106-998e-48e3-8e3c-78549edb9631 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ToolRL: Reward is All Tool Learning Needs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.104200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.104200Z digest=sha256:b792bb11ee7821e129e60999d0b1e8d91976b3dc15a90ac5ffb7feccacd5d993

Observation 774b04e1-77b9-4522-a558-b8866e1da4a4 · outbound

This paper cites Plummer, Liwei Wang, Cristina Cervantes, Juan C.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Plummer, Liwei Wang, Cristina Cervantes, Juan C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.339740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:25.011371Z digest=sha256:f47532f1648aab1911dec76f17ffc3be8bc4010ad7b42ac164df66e74a341afc

Observation 9c791ef3-4db7-4304-a588-3e1aef87e671 · outbound

This paper cites Autogpt: An autonomous gpt-4 experiment.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Autogpt: An autonomous gpt-4 experiment

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.185113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:25.302039Z digest=sha256:b50d4b6a933032fd9f640dac0ba094b7007ed3180e99c36e7a5816ced4856401

Observation 6aa3af0b-9d82-40a7-9d69-c5b52652698d · outbound

This paper cites StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.216062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.216062Z digest=sha256:acc2cf4b8e63f6b86fd5fcfade44f5c15c451b45226bb3a106df7a14b02a0f9e

Observation ee478595-f1e9-4889-9292-680e2f3b3052 · outbound

This paper cites Partially observable markov decision processes.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Partially observable markov decision processes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.490284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.490284Z digest=sha256:cff04ed2da52442da6e18dc5e77f50184fa73fae1dfca91df5e52e3fb1766245

Observation 9a8f005a-2bfa-4f72-b3bb-f717f73a7e5f · outbound

This paper cites Towards vqa models that can read, 2019.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Towards vqa models that can read, 2019

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.390603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.390603Z digest=sha256:ff71bab5d60d5825642781215ff47fb0b8dc942e1c704b0035afacc77c117fc9

Observation 1d533068-f779-403b-b4fd-e6fedf97e8dd · outbound

This paper cites Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Phrasecut: Language grounding in images by text-based mask segmentation.ECCV, 2020

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:28.038738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:25.671645Z digest=sha256:21ccd427698494bf6c698b8b144b1dfa52d10a907a756d6e129bd5ba537cb4e6

Observation 06cac098-554d-4f85-97d7-084e3f74b64f · outbound

This paper cites AdaPlanner: Adaptive Planning from Feedback with Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents AdaPlanner: Adaptive Planning from Feedback with Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.568370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.568370Z digest=sha256:bb6de73da1c4684529c96056349bf4c0923a7e2af2acac861fe4f612b8074ccf

Observation a70e694e-7803-4b19-a508-f58b603a9a32 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.853689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.853689Z digest=sha256:d4ef3ded19e20864d1b716153aee28d5a107ccac3fa55058fb90bb9ca786d077

Observation a684f4db-af1e-4cbd-9758-0b86697091d1 · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.770961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.770961Z digest=sha256:7e09555fede82eefe773ecd6c62ae183c8dea78f6c2a5eaa30c22d99a6f880bc

Observation 7c0b0e93-0804-4c73-af47-4ef6676fbd90 · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.066956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.066956Z digest=sha256:f360c9190db126999f9c6517a91a630b9f7c0295fd3f3b5a9c22b05084d9f4f7

Observation aeaf88a7-c605-4ead-ae1a-6f075e29978e · outbound

This paper cites DeepSeek-V3 Technical Report.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:25.977120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:25.977120Z digest=sha256:c6098d245957793d582e4840e143e94ab8f6b94f2cad5bf7a802f29027f9e08b

Observation 3be8dd6d-a812-4ac5-920a-78dbcd120b33 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.288484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.288484Z digest=sha256:246d59de1655719c1d6e10f2c62b740329c1a344d85db4900296bd339be33613

Observation 1bd161c2-6d3c-4a36-99b9-a4db26cae643 · outbound

This paper cites An illusion of progress? assessing the current state of web agents, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents An illusion of progress? assessing the current state of web agents, 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.859356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:26.170273Z digest=sha256:2dc8874b0c88714dde51a182d51624e88de800bb47370800ca3fa53577b6bd4f

Observation ceaa361b-e936-4422-af1b-2d0fb9da548e · outbound

This paper cites Berg, and Tamara L.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Berg, and Tamara L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.494034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:26.499271Z digest=sha256:313032ee513aaa68085fadffe3326a8853dc7d8938c9f514817e22dbe03db848

Observation 24ae8380-6409-4b5b-8208-233f31bdc736 · outbound

This paper cites Survey on evaluation of llm-based agents, 2025.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Survey on evaluation of llm-based agents, 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:16:27.702921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:26.402569Z digest=sha256:4b67496a4b9fd7e82eddb3bd1c563816a3b8b2198bafd927f15d1ea7405cf32f

Observation 162ccc53-e690-435f-88a0-cef2f6fc6705 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.677705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.677705Z digest=sha256:760a08f381a7a7df9036c20b46a59dad7b325371bb7938d03cb0f6601fcaee91

Observation 7a051344-355f-4ef5-83e4-820e08d2ce20 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.604788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.604788Z digest=sha256:f9270fd3ecbf5cd65eeba49134624c21aba82d3ff9142e42b3d156ee15fc4ca3

Observation 5e1f3d13-c12b-43b4-ae24-d0ae1f1b2938 · outbound

This paper cites PubLayNet: largest dataset ever for document layout analysis.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents PubLayNet: largest dataset ever for document layout analysis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.842977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.842977Z digest=sha256:8bdcee8ff658424815ae9672c5b7055ec7bcf672edc9729f19a9d076d84e032c

Observation 7932a883-2bfe-4888-803c-20a7cea739a8 · outbound

This paper cites Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.754466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.754466Z digest=sha256:016423ea414553ee9c84e27ac37a61494873d2432a569e17b71ac6fa6a26882b

Observation b0ebfb9b-12ba-458b-a62c-90823959b5cb · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:27.016316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:27.016316Z digest=sha256:d7643af4ee58365d4752f2104dd80aa0da252200ec7d020c13bfba33c6a88d05

Observation 6a82974c-ce36-409f-a4b8-61390314f9de · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:26.935337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:26.935337Z digest=sha256:62a313a7c617d401b2d510df8b003d6e99a2453ef00fe897049a4b6855186d79

Observation 3a877dd6-c95b-48b8-94c7-068352c75590 · outbound

This paper cites an unresolved cited work.

Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T12:16:29.739225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:16:22.800470Z digest=sha256:269c9c658bea0fa0810d26192cb609763cade6b23ec3e7223ecc8ace58208a7d

Pith citing papers

Observation dc9c335b-065f-405c-a134-7afa15c65c37 · inbound

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges cites this paper.

Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 167

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:42:22.039877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T03:42:10.703369Z digest=sha256:ae01341f515dbb8342f2b1642f0d1abab458bb212b4397c402487a946ee9ca8c

Observation 7f2c9b87-98f2-4dbd-b6e8-d8c6495768d7 · inbound

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers cites this paper.

COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:23:57.983899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T03:23:03.198463Z digest=sha256:ed5b5a7d4ccdcb5109d8e4f090a1d2a8ddb56f36d6808ef6d77d53b30fa68928

Observation b2f85212-f096-43df-936e-dca4295716bc · inbound

HLL: Can Agents Cross Humanity's Last Line of Verification? cites this paper.

HLL: Can Agents Cross Humanity's Last Line of Verification? Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.583281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T14:57:57.218669Z digest=sha256:189d6a91530f672ce244f4d91a45ad371b1f7eb66e472550661d9c12749631ed

Observation 1e9173ce-1855-42c5-9deb-ca06122ddbce · inbound

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents cites this paper.

Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T14:45:56.506288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:45:56.506288Z digest=sha256:75eedc138505e102cb8c4962bceae5af2a02c19db152004077ee2b930edf4b22