Pith. sign in

Paper Citation Record · LEDGER

Benchmarking LLM Judges for Mobile Agent Evaluation

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.11434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11434 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:32.198253Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 459565ae-63e7-4dd6-a812-ad3413836f2f · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.911176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.911176Z digest=sha256:978ecef6d432af1b33358c040217463f827f98f7a296de232ff9bfe6db4fa431

Observation 4c37776f-8270-447b-8722-155012e39f7b · outbound

This paper cites A Survey on LLM-as-a-Judge.

Benchmarking LLM Judges for Mobile Agent Evaluation A Survey on LLM-as-a-Judge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.917445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.917445Z digest=sha256:24f98709c5335a2d413891a2abfda180939a042fbe5dddec339fade0e1525181

Observation 859ba63b-7930-4f89-a197-ccde841f800a · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Benchmarking LLM Judges for Mobile Agent Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.923269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.923269Z digest=sha256:91a42ae56f6ea9b8f30c17ddb702ef8fa41f6b8cab145e193b280665ad629616

Observation 22ce04ac-f32a-4eea-8542-f17a247c9d6e · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation Agent-as-a-Judge: Evaluate Agents with Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.928785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.928785Z digest=sha256:b6f80d5a156256b4a752851d909a4955b6621a9c3520a5aaaad60956f603db8c

Observation 4159f734-2f57-418f-9a2e-69c61e1657ef · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Benchmarking LLM Judges for Mobile Agent Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.934258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.934258Z digest=sha256:8bcdee8c3128ed33b30392db010a9ade922e7c84dac22d2343af56732b0b47ac

Observation d1fa04e0-638a-4413-8c14-96337f107169 · outbound

This paper cites arXiv preprint arXiv:2502.01534 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2502.01534 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.939821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.939821Z digest=sha256:a84cf0e6c8ad37f7105b204e1930d49ba0789be84a9d05d84a9d220a3ae7033e

Observation 59a8ad62-7e33-48fa-b8b8-bdeb29d4b960 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Benchmarking LLM Judges for Mobile Agent Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.945269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.945269Z digest=sha256:98e27290b60a3811ddaa63bbb55e2e47c5c9de50140e481caf32b0638f462dd0

Observation 04cf6c93-592c-4d2a-afaa-6fbdbd114091 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Benchmarking LLM Judges for Mobile Agent Evaluation LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.951140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.951140Z digest=sha256:d87732e543878fe3465733e8c4672de2f694fd53b4d913ce4b9e0442de8e8126

Observation 53e2ec39-3c31-494c-baaa-6429e480fa76 · outbound

This paper cites arXiv preprint arXiv:2504.08942 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.08942 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.956961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.956961Z digest=sha256:76dc850f77cf5587baadcd1560649c67778fbaafb5eeee424c13974a97df6b69

Observation 032d1b18-e46a-41d3-90a3-a4e053a48ce1 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation Autonomous Evaluation and Refinement of Digital Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.962001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.962001Z digest=sha256:549d5feb7f51562df6d843c25e548e2c9cef19be081a9dbe7e4310b376b00ef4

Observation 098aef06-4c47-4b4e-82b4-11f191ecb904 · outbound

This paper cites arXiv preprint arXiv:2503.02403 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2503.02403 , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.967653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.967653Z digest=sha256:fd053c99b5dc0fc7a83cd34c3c8a191c1d0df956197109be6e24755ea8264691

Observation 416e280a-2822-4573-b6ef-49d966f55767 · outbound

This paper cites arXiv preprint arXiv:2504.01382 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.01382 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.972605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.972605Z digest=sha256:48e3b60a168b2790ca3a260a6470d09b37cf8d185d5bf98be833756aac88615d

Observation 785bddc2-7648-4494-be09-4673ab7c0821 · outbound

This paper cites Agentic Reward Modeling: Verifying.

Benchmarking LLM Judges for Mobile Agent Evaluation Agentic Reward Modeling: Verifying

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.355751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.978148Z digest=sha256:d1d295ab0565041e65b596a37aa22d21ffb38802880e1f48caba5365e0685dbf

Observation 8963812d-7c3d-4d88-8712-ca0ae5ca3e97 · outbound

This paper cites ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration.

Benchmarking LLM Judges for Mobile Agent Evaluation ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.983082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.983082Z digest=sha256:2c6b2179b1543bfbdcd06a1022bff8f39f42238e5ac319b81d9e40ee8ee779a2

Observation 84019af1-727f-40e6-a9ce-30c178af32cc · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.988574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.988574Z digest=sha256:81a10394eb1d86e3ceea0362e78806e91bfe61e3a0d4611623130af7222ae668

Observation 51f932d1-9049-4148-b46e-9b8eb2ec1d99 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.326938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.993527Z digest=sha256:d125ac8af6b377233b3e6c4706e8b36e168ec70c33c5c84abdc0d5dececcf8bd

Observation f2883150-765b-44d4-bbb6-7f7d9e876b27 · outbound

This paper cites 2025 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.310381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.998716Z digest=sha256:cd81163ffd0c3662a629af94541c2857d551fff630386f6fb15b7b3a6a94cd5e

Observation 8ce83736-aed8-46d5-b7aa-a39987ff5fee · outbound

This paper cites Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.295066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.003906Z digest=sha256:80b9e595d090b7fddd2ebc998f097e4545f6083e002d50f73462e7c91b8cdc68

Observation c76913cf-4f77-4859-8e4d-d017b199c749 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.008836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.008836Z digest=sha256:62f52db730a6bdee56bcfe2b6ab181b5da23211af61bbd573e2665a899f36ac9

Observation 20a63972-cfe0-4996-ad35-5a3959ff5151 · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Benchmarking LLM Judges for Mobile Agent Evaluation Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.013673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.013673Z digest=sha256:af695a1b9a1a23aaa7a2bff456e09315593a98a16ee78fa00134bdf557bd83d2

Observation 0734db5b-9877-426b-a87b-8099740f23e8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.018875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.018875Z digest=sha256:5a7ffab0285834bb67c256681342e5f4374c4b7fbff89707a3f9fbee8fda6ded

Observation 856bbb72-5629-41d5-9e79-8e999dc5a31e · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Twelfth International Conference on Learning Representations , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.023625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.023625Z digest=sha256:2f1187ec9c992e712d3695f0ffb823c11c289accae07ade20df0c3c77b0b2663

Observation 40982bc3-e1a1-48a3-a05a-b807dd17939d · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.028406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.028406Z digest=sha256:6ce5dd10b335304f1b2935496a9913db3831d5c9b946655be944d4ea4bb357c6

Observation c06a0372-89c0-4a33-8383-574f9f29d717 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.033137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.033137Z digest=sha256:2a48f5d05c165ded6e1004a13afab48d6c880acea2b49daf8167c3431cd449d1

Observation 3edaa242-13ea-49b3-ba36-5ac3246d2f18 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.038168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.038168Z digest=sha256:db04703160ebeb8a842cc48da6443d260d17637090f754aaff6267a119b428dc

Observation e6b39dcc-4d15-4b58-b11c-f377d615c79f · outbound

This paper cites AgentBench: Evaluating.

Benchmarking LLM Judges for Mobile Agent Evaluation AgentBench: Evaluating

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.042790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.042790Z digest=sha256:cf82d58fb7378b9c4533175a981f882cc013fd360a1086d9309a0091bd91419e

Observation 44b90881-c4fa-42fd-9c9a-737df97b0901 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.047694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.047694Z digest=sha256:16f70c4812fd5d1b68b37ead606ee2be82fec3bf5ffbd66732ccb1ae53582e1a

Observation 174e94ea-8eff-427a-8fd1-b262c726d22e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.052364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.052364Z digest=sha256:b35a18db4b0b51151ada917a87a0ed203ec222b86ea7693aadbe633dc95474bf

Observation 0d672b7b-1725-44eb-bf82-ced1257f74e2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.057692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.057692Z digest=sha256:04f024ee1929f18428579c5d9fdd35f42727386b68deef64d7fbe24da2725c1b

Observation 172e35b6-9782-4df1-82a2-1f2450f4ab0a · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.062580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.062580Z digest=sha256:14a651f11bfd4936b6bc818a6243d5cb6c456e209e6f63c44d5fcf312e6b913f

Observation 6cbe5f76-0556-41be-a8aa-af24c2fe83a4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Benchmarking LLM Judges for Mobile Agent Evaluation Constitutional AI: Harmlessness from AI Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.067446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.067446Z digest=sha256:79b50753120e40fa824649b2ecfc73e92df49d813363ecfdcb6a43a9d39f7a19

Observation e9bd76ce-9bd4-40ed-b964-3c170892c2a6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.073219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.073219Z digest=sha256:76fa27df335992462d79db837b4f89c3eb9e54ecdbb54f736e80c2ef18fd0d06

Observation 03355267-5ae0-4dfa-b3dc-df89f733b8ed · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.079033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.079033Z digest=sha256:2be23e4ab61a5615bf9d1568d2127ada3715730070337c3016618530f5029541

Observation 296ab333-57e0-4942-948f-8679fa6474f2 · outbound

This paper cites arXiv preprint arXiv:2509.18119 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2509.18119 , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.083855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.083855Z digest=sha256:5972f2e4c81f3ceecedd62ffa08e9cbaa9e3e63db4730000df964dba5a94c561

Observation 0f56b061-3d55-465d-a1bf-fc66b6dc4e69 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Benchmarking LLM Judges for Mobile Agent Evaluation MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.088344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.088344Z digest=sha256:1daf7e051ede7d213c47097004aee7f4368c82889610ae92f4ee70117723aaaa

Observation 65c1156f-ba41-4315-a70f-56cb60bbc993 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.093292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.093292Z digest=sha256:0c1fb024099d694b88f4e5eb9326581e5901543b220499b3bbeabff9255227e6

Observation 77aeb3c1-2a88-4476-9888-5e1845537f1c · outbound

This paper cites International Conference on Machine Learning , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation International Conference on Machine Learning , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.098421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.098421Z digest=sha256:9640614d8d1bbe1b4f4aa7e90a36bba996a43970943f4a17f0c0880eb906c7d0

Observation fdc80167-d9ce-44c6-bdd7-9fc0bea76df0 · outbound

This paper cites arXiv preprint arXiv:2409.15922 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2409.15922 , year=

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:16:32.473314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.104134Z digest=sha256:6244d8567c037eee7d0fba99f917ba9adeb6cf0ba29c012f14fd53a39045856f

Observation 596d336c-524c-4359-956c-5136dfa0cb19 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Eleventh International Conference on Learning Representations , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.109139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.109139Z digest=sha256:02b12d94b734a947a8748d5832d55b78c75f799877fc89f57233ce49a8df666d

Observation 8ab4694e-adef-4b5b-a7e2-ad6578e1705d · outbound

This paper cites 2023 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.113664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.113664Z digest=sha256:d97894ae161f5074d7fa4824d6b135ec065da14142651d33995d6e3b1c78f7a2

Observation 42469195-bcff-40ab-b82e-fcd6fd95bd22 · outbound

This paper cites 2023 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.120720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.120720Z digest=sha256:20bccefc5eb960724abf5952aca8e70be1a6faa0aa88f5d3325cfcd781ce31b4

Observation 467705c5-13bc-4cf1-bce6-df9e156c3a3c · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.125755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.125755Z digest=sha256:b39648aed59253e9088e57a21d733be2bc4e1ebf2329968ea90a2272bf41bf73

Observation da50c88c-d8e9-4378-ae52-fba1444e3721 · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.130803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.130803Z digest=sha256:9338759248febb76bf9ac604fdc2281f51b402c380dff239998b45b9a52a9a71

Observation 813b31c9-9785-484a-896d-4cc80db9f165 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Benchmarking LLM Judges for Mobile Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.135869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.135869Z digest=sha256:05c8d354080f4f37d7410cad409e54f0e2ba1e9c11272abb9c8fcd879763e3e3

Observation db4c014c-b9a1-4712-8858-0f5ad09dbf9d · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

Benchmarking LLM Judges for Mobile Agent Evaluation NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.141202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.141202Z digest=sha256:fecdf4664aba444dbc70d8b381b6d3c8c94a103b21f1b6201c81bcef4def7c11

Observation c169458d-480d-4331-b420-f2e2239d0b6c · outbound

This paper cites 2025 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.146221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.146221Z digest=sha256:0e0adceb9e7887306e48b34c2d6d1402057d58993235370d7d98e8d9fb431246

Observation 65c21e3b-1e6f-475c-a3cf-39d8ca993a1e · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Benchmarking LLM Judges for Mobile Agent Evaluation ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.151042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.151042Z digest=sha256:3b6f17f51f6dc11f4a906e406698586f30ed8e672b347c77ee5a324663d46bdd

Observation c2ca3d5b-87f6-4e61-9d75-a6c634fbc8c9 · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking LLM Judges for Mobile Agent Evaluation The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.156301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.156301Z digest=sha256:94f924bdd58d60593550aecdeaf5a499b6ef93353eac719fcc5f16f4d73c1315

Observation 9d0df293-bc95-447d-9e23-29eba29c88e9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Benchmarking LLM Judges for Mobile Agent Evaluation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.161977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.161977Z digest=sha256:15f0230be5e16cf1b4f1d258e739a5aea1377fe9eed4dc774f2146fb17eb1fac

Observation 2fa48a6d-d961-44df-b778-5fc06d7491e0 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.063532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.167383Z digest=sha256:6530494b4ff3ce61c783a4a6fb72534aeb8a130787d8e720e9ab57ebb0147e22

Observation ecc1b69f-87eb-4487-9980-26640e937fbd · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.046366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.173421Z digest=sha256:3eb13132447a6f9e8ddbf124ed29fe330fd03fb14664a54071d8992d76b3eea8

Observation f791c250-faf3-4e16-b380-307aca291450 · outbound

This paper cites 2025 , note=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.027992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.178476Z digest=sha256:194f92444b5a23dedc6a83d4bd70a2994b749f3a1febfacba001ecf7d7cdeada

Observation eb6527f1-0fa1-40b0-8443-0b67c865ece4 · outbound

This paper cites 2025 , note=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.183383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.183383Z digest=sha256:ed546d4af732d1511a6d5bf92bbc716bd31890f5834be0ad7f4e587c2ca6a36d

Observation 248cf1fe-a7ce-41c4-a806-0b0cc83d4ad2 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Benchmarking LLM Judges for Mobile Agent Evaluation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.188278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.188278Z digest=sha256:7ef35e0a87c2b11527d1a2c4d6ae16a2485a25f5a7333cef827023b8a044e16f

Observation 98d9f3cb-4d8f-45a5-83ab-d70e1bdd8a69 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.001354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.193184Z digest=sha256:c2a1f2248c598b0087a58ee16bcfa8243f88a7305b5d39aea2b4ec994c20d23c

Observation 347f05eb-0442-49f7-9010-a9d9604287db · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:32.983736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.198253Z digest=sha256:56be41840ca8a0e63ecc41d2b0d56a15890228403eacd770f91b53e5a0c0c0f4

Pith citing papers

No inbound Pith citation observations are available.