Pith. sign in

Paper Citation Record · LEDGER

Benchmarking LLM Judges for Mobile Agent Evaluation

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.11434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11434 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:16:32.198253Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 459565ae-63e7-4dd6-a812-ad3413836f2f · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.911176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.911176Z digest=sha256:76828d18ca88f69c805f53f8c47a959f525c164ab6b652cc5a92d4c62fe1fd1f

Observation 4c37776f-8270-447b-8722-155012e39f7b · outbound

This paper cites A Survey on LLM-as-a-Judge.

Benchmarking LLM Judges for Mobile Agent Evaluation A Survey on LLM-as-a-Judge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.917445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.917445Z digest=sha256:84cb6d31553fe77069b3f5a6a9c3a07873039db0f9a2ca711266f3c67e32fa91

Observation 859ba63b-7930-4f89-a197-ccde841f800a · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Benchmarking LLM Judges for Mobile Agent Evaluation LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.923269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.923269Z digest=sha256:806f0f5edb55f9fa1876918bc78999cb49a00f71c09996c0af3cbcc55b17ed40

Observation 22ce04ac-f32a-4eea-8542-f17a247c9d6e · outbound

This paper cites Agent-as-a-Judge: Evaluate Agents with Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation Agent-as-a-Judge: Evaluate Agents with Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.928785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.928785Z digest=sha256:302ea90612766a07b4de905a08d57e3c5f30186098cbc7427f037469782d6de8

Observation 4159f734-2f57-418f-9a2e-69c61e1657ef · outbound

This paper cites Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement.

Benchmarking LLM Judges for Mobile Agent Evaluation Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.934258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.934258Z digest=sha256:a41c57cd5105a8c147c615b3bdcc788f6a9e5702f1c5764572acd77b39deee47

Observation d1fa04e0-638a-4413-8c14-96337f107169 · outbound

This paper cites arXiv preprint arXiv:2502.01534 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2502.01534 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.939821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.939821Z digest=sha256:dbb3506afbc6d13366da3eace7b663a684ecc3972d33bc3ef27dcf5964296f31

Observation 59a8ad62-7e33-48fa-b8b8-bdeb29d4b960 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Benchmarking LLM Judges for Mobile Agent Evaluation JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.945269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.945269Z digest=sha256:05616553d5f41a5b239f129b89ce14d2ad7b4eb34112415b3da55027d81adfa6

Observation 04cf6c93-592c-4d2a-afaa-6fbdbd114091 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Benchmarking LLM Judges for Mobile Agent Evaluation LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.951140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.951140Z digest=sha256:1f67a91ce20eda950bfb238b2d7931e222481240de2d0eb5988262421e44e76d

Observation 53e2ec39-3c31-494c-baaa-6429e480fa76 · outbound

This paper cites arXiv preprint arXiv:2504.08942 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.08942 , year=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.956961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.956961Z digest=sha256:205c851ef7ecdce1da1dfa10c462757d138a92bae0eebee54f207e469310ef0d

Observation 032d1b18-e46a-41d3-90a3-a4e053a48ce1 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation Autonomous Evaluation and Refinement of Digital Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.962001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.962001Z digest=sha256:ca0af12200de90f4c990adf143a5918f040b2b3809a523b88a8da3f760a99edf

Observation 098aef06-4c47-4b4e-82b4-11f191ecb904 · outbound

This paper cites arXiv preprint arXiv:2503.02403 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2503.02403 , year=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.967653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.967653Z digest=sha256:778aa07b35d883da82895811fac2d6cf12a962aa1b82431abb4383a86652c5a1

Observation 416e280a-2822-4573-b6ef-49d966f55767 · outbound

This paper cites arXiv preprint arXiv:2504.01382 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2504.01382 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.972605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.972605Z digest=sha256:f5b485f44a9d3832efc07e6fa7c64c6226db02a7fcba1d9497e54796c54c0595

Observation 785bddc2-7648-4494-be09-4673ab7c0821 · outbound

This paper cites Agentic Reward Modeling: Verifying.

Benchmarking LLM Judges for Mobile Agent Evaluation Agentic Reward Modeling: Verifying

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.355751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.978148Z digest=sha256:007f1bf1f77ed399e183227195b94147e5844fe1105b5925668121a31627a1e5

Observation 8963812d-7c3d-4d88-8712-ca0ae5ca3e97 · outbound

This paper cites ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration.

Benchmarking LLM Judges for Mobile Agent Evaluation ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.983082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.983082Z digest=sha256:132871a7c7c44a3af368e61245b27d0bc3c55985eeefed1481f0abbe75ad4ba1

Observation 84019af1-727f-40e6-a9ce-30c178af32cc · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:31.988574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:31.988574Z digest=sha256:2d9bebade37655462951c623a3d281bcd940081d8f63bd19ac4d6ac4a0ff4f84

Observation 51f932d1-9049-4148-b46e-9b8eb2ec1d99 · outbound

This paper cites The Thirteenth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Thirteenth International Conference on Learning Representations , year=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.326938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.993527Z digest=sha256:ff991d0851063fbe8696c3d4b0230a495820e8fc7464a88f75d4f4daa2d3bf2b

Observation f2883150-765b-44d4-bbb6-7f7d9e876b27 · outbound

This paper cites 2025 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.310381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:31.998716Z digest=sha256:698445a32a7900652f6a8290df332d4d096d100dff296f4b15bdc4a37dff374d

Observation 8ce83736-aed8-46d5-b7aa-a39987ff5fee · outbound

This paper cites Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 30th ACM SIGKDD conference on knowledge discovery and data mining , pages=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.295066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.003906Z digest=sha256:27fe74ab46490c673a784fee27780159ee82b17a9ca7891bd60ec16bfd84536f

Observation c76913cf-4f77-4859-8e4d-d017b199c749 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.008836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.008836Z digest=sha256:4b66654a6ca427ba4fa47e4fae10016cff7d49a0a7750dbc98b5c64f42276300

Observation 20a63972-cfe0-4996-ad35-5a3959ff5151 · outbound

This paper cites Benchmarking Mobile Device Control Agents across Diverse Configurations.

Benchmarking LLM Judges for Mobile Agent Evaluation Benchmarking Mobile Device Control Agents across Diverse Configurations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.013673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.013673Z digest=sha256:9a13b5060a41cb34d159e64b143a279ae696a06a9ddb7e8ca46a4527e75a530c

Observation 0734db5b-9877-426b-a87b-8099740f23e8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.018875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.018875Z digest=sha256:9f4b026f4f7f69c394e34474111ce03f68855362b4d5bf0b6c236fea26f9955d

Observation 856bbb72-5629-41d5-9e79-8e999dc5a31e · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Twelfth International Conference on Learning Representations , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.023625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.023625Z digest=sha256:c342c05083bfcc21591fd766d1e8266472e01ec914c5b53e568c9f0ba704bf99

Observation 40982bc3-e1a1-48a3-a05a-b807dd17939d · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.028406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.028406Z digest=sha256:f22fa029e6bb0073ee505a282bb2630a4c26f6150464fd399c230abd2003626a

Observation c06a0372-89c0-4a33-8383-574f9f29d717 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.033137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.033137Z digest=sha256:513ce0e042328f158c23d645b10cb996981105d152d002316b5f762a3ddabdf8

Observation 3edaa242-13ea-49b3-ba36-5ac3246d2f18 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.038168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.038168Z digest=sha256:dd69d60ba4cc63a61390c8106a02e2672320afaa60196663378984f4edeac25f

Observation e6b39dcc-4d15-4b58-b11c-f377d615c79f · outbound

This paper cites AgentBench: Evaluating.

Benchmarking LLM Judges for Mobile Agent Evaluation AgentBench: Evaluating

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.042790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.042790Z digest=sha256:825339cf680943a623a65cfd4634ddba8e0c48d7c7cb0c9ef69e9ca16da05b93

Observation 44b90881-c4fa-42fd-9c9a-737df97b0901 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.047694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.047694Z digest=sha256:d6ffa053b18cfde37608f588b85ff6b5b369951ba028fb9709a2702973714a04

Observation 174e94ea-8eff-427a-8fd1-b262c726d22e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.052364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.052364Z digest=sha256:c3226968435c4f1f664937c830c8c19ad85789d90e1c1b6eff06acab5ba3a9e1

Observation 0d672b7b-1725-44eb-bf82-ced1257f74e2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.057692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.057692Z digest=sha256:8ef2ecc5abeb7c37b7b8ff86bfcaa1f6d87cc0b78389dd1ab925caeb0c1847be

Observation 172e35b6-9782-4df1-82a2-1f2450f4ab0a · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.062580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.062580Z digest=sha256:9e62866fd38b31d14549a25246d0b8e82827188e9707e83c0ac006daa29932ad

Observation 6cbe5f76-0556-41be-a8aa-af24c2fe83a4 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Benchmarking LLM Judges for Mobile Agent Evaluation Constitutional AI: Harmlessness from AI Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.067446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.067446Z digest=sha256:54658901269ae28c825eb4e8fe4614de4a5a8101a8bfa1a09ca31c04fcbc63a6

Observation e9bd76ce-9bd4-40ed-b964-3c170892c2a6 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in Neural Information Processing Systems , volume=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.073219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.073219Z digest=sha256:e2cc274f52f301018bbbf697a95cad3b3cf70f916668116375ffa335eb45362f

Observation 03355267-5ae0-4dfa-b3dc-df89f733b8ed · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.079033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.079033Z digest=sha256:03fd36cc3b848f9a7b23b2ac6f221be3b1d1fd8f8b16806ba7afd35a3ce47bc9

Observation 296ab333-57e0-4942-948f-8679fa6474f2 · outbound

This paper cites arXiv preprint arXiv:2509.18119 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2509.18119 , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.083855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.083855Z digest=sha256:10690cefff643d37f5c3fee686aa571a1278a8f2be32bc50d45c4c5a30ec0e14

Observation 0f56b061-3d55-465d-a1bf-fc66b6dc4e69 · outbound

This paper cites MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment.

Benchmarking LLM Judges for Mobile Agent Evaluation MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.088344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.088344Z digest=sha256:330006548271bf9dfa6c6d987f3d36256c921bf89341bc275d1db57722c3f10e

Observation 65c1156f-ba41-4315-a70f-56cb60bbc993 · outbound

This paper cites Findings of the Association for Computational Linguistics: NAACL 2025 , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation Findings of the Association for Computational Linguistics: NAACL 2025 , pages=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.093292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.093292Z digest=sha256:1df4af78f8e719ec38f83b02278ed19ccaaaebfe867a3cd9ebc9cf0ee91c2779

Observation 77aeb3c1-2a88-4476-9888-5e1845537f1c · outbound

This paper cites International Conference on Machine Learning , pages=.

Benchmarking LLM Judges for Mobile Agent Evaluation International Conference on Machine Learning , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.098421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.098421Z digest=sha256:edcb5c2cce2547045bf1236a07c2ee77b6289c1813bf8e0f9cc21a7100cd6322

Observation fdc80167-d9ce-44c6-bdd7-9fc0bea76df0 · outbound

This paper cites arXiv preprint arXiv:2409.15922 , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation arXiv preprint arXiv:2409.15922 , year=

Reference 38

Resolution
verified exact
raw_fallback, observed 2026-08-15T14:16:32.473314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.104134Z digest=sha256:dfddcad23d2bd4aeffd31ef866a37dd7fac9e8e2627be47660bfe814a60f25d3

Observation 596d336c-524c-4359-956c-5136dfa0cb19 · outbound

This paper cites The Eleventh International Conference on Learning Representations , year=.

Benchmarking LLM Judges for Mobile Agent Evaluation The Eleventh International Conference on Learning Representations , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.109139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.109139Z digest=sha256:30f2138c09c72fa0bc9cef13ca5c8e8ee3578ceb84026db340f9df422a33611e

Observation 8ab4694e-adef-4b5b-a7e2-ad6578e1705d · outbound

This paper cites 2023 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.113664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.113664Z digest=sha256:8f220d92863734e15bbf5fb1cf449e32f0656f32cdfcbacd57fe7659e2b221d3

Observation 42469195-bcff-40ab-b82e-fcd6fd95bd22 · outbound

This paper cites 2023 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2023 , eprint=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.120720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.120720Z digest=sha256:af9fe64e2e1501d2ef69df4d5f18a2b968285214106d578180f4f12f15564655

Observation 467705c5-13bc-4cf1-bce6-df9e156c3a3c · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Benchmarking LLM Judges for Mobile Agent Evaluation UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.125755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.125755Z digest=sha256:3ede91095256fff9cd92b666e14b64725a5ae793872645dc46c7c5abd4dff6c6

Observation da50c88c-d8e9-4378-ae52-fba1444e3721 · outbound

This paper cites Advances in neural information processing systems , volume=.

Benchmarking LLM Judges for Mobile Agent Evaluation Advances in neural information processing systems , volume=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.130803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.130803Z digest=sha256:2b0ef9260c143cb8106b0ccff1166b415a3999d2cf777555ee2eaa6dd1b43ca6

Observation 813b31c9-9785-484a-896d-4cc80db9f165 · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

Benchmarking LLM Judges for Mobile Agent Evaluation WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.135869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.135869Z digest=sha256:ec592b13cf981c480a5d0dbf08334a8ca7ce2aca92612646230c8e165c90f883

Observation db4c014c-b9a1-4712-8858-0f5ad09dbf9d · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

Benchmarking LLM Judges for Mobile Agent Evaluation NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.141202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.141202Z digest=sha256:3603a23678083faa81e2d8202151ccabed8fb62d85abb2285a8fd29dba531acb

Observation c169458d-480d-4331-b420-f2e2239d0b6c · outbound

This paper cites 2025 , eprint=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , eprint=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.146221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.146221Z digest=sha256:23fbd472d2bf4597c170c5d0d78d941f6c130eeabd164fa54cd6419072ba21bd

Observation 65c21e3b-1e6f-475c-a3cf-39d8ca993a1e · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Benchmarking LLM Judges for Mobile Agent Evaluation ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.151042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.151042Z digest=sha256:3adc112c2514544d0226e6980dd9418fd70f62568295647f783c0638758bbacc

Observation c2ca3d5b-87f6-4e61-9d75-a6c634fbc8c9 · outbound

This paper cites The Llama 3 Herd of Models.

Benchmarking LLM Judges for Mobile Agent Evaluation The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.156301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.156301Z digest=sha256:f97d7ec3620d2426e43dc9802afa6d07971b1d3fb777e145fa98492977580a7b

Observation 9d0df293-bc95-447d-9e23-29eba29c88e9 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Benchmarking LLM Judges for Mobile Agent Evaluation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.161977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.161977Z digest=sha256:d9a3aba2d366bd21ce14e8c27ee3b3f3f153e953814012f7596144d553b7848b

Observation 2fa48a6d-d961-44df-b778-5fc06d7491e0 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.063532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.167383Z digest=sha256:af8f7e35fc67ccb33a0f7784073be598111620e532a3339f05300808da136669

Observation ecc1b69f-87eb-4487-9980-26640e937fbd · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.046366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.173421Z digest=sha256:c8523960b0d5d3dc928359be1c5d12ca627fa63c5c685df21b3aa6db1485f670

Observation f791c250-faf3-4e16-b380-307aca291450 · outbound

This paper cites 2025 , note=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:16:33.027992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.178476Z digest=sha256:fd64138b55159f9e52c59e24f96de61c417818370bf48c5545f56cdaa3849c82

Observation eb6527f1-0fa1-40b0-8443-0b67c865ece4 · outbound

This paper cites 2025 , note=.

Benchmarking LLM Judges for Mobile Agent Evaluation 2025 , note=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.183383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.183383Z digest=sha256:8ef63f4d14ca9d1c6641f129845863f5363f0cffc12be60e1037e3f4aa7fe747

Observation 248cf1fe-a7ce-41c4-a806-0b0cc83d4ad2 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Benchmarking LLM Judges for Mobile Agent Evaluation GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:16:32.188278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:16:32.188278Z digest=sha256:ba17c6b276b9adc9d2e5d7e463d139093dd1fa937237c8c5702f4a5c750d9207

Observation 98d9f3cb-4d8f-45a5-83ab-d70e1bdd8a69 · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:33.001354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.193184Z digest=sha256:edb118832ea4bd16268e76dc3731c2a6dde1ae03d636d2d341d3c274aa4d7c57

Observation 347f05eb-0442-49f7-9010-a9d9604287db · outbound

This paper cites an unresolved cited work.

Benchmarking LLM Judges for Mobile Agent Evaluation Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T14:16:32.983736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T14:16:32.198253Z digest=sha256:281f8a76d8d33c7549006ae60513f678b8a6b6858b450bf070860bbbf99f439d

Pith citing papers

No inbound Pith citation observations are available.