Pith. sign in

Paper Citation Record · LEDGER

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

As of 13 August 2026, this Paper Citation Record lists 100 of 218 outbound references and 0 inbound Pith citation observations for arXiv:2608.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10692 v1

Coverage vector

measured 100 of 218 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.257857Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 218 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3d2b53-cf36-4553-99ae-37d9bec5e39a · outbound

This paper cites Langley , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Langley , title =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.624053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.624053Z digest=sha256:d3301eb452f2d2a676b5d376e736a578825613fdbc4cac8c0df2244623d20ca3

Observation 2e475699-9ac0-4f50-ab64-29a8dff91f87 · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AutoGLM: Autonomous Foundation Agents for GUIs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.630665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.630665Z digest=sha256:809b9027971883e7edc1629dc948ab2b7140adea277aa4d449827587ca4bab3c

Observation 2cbf283e-58e2-4f43-a012-a5ff84ee17f4 · outbound

This paper cites Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =

Reference 3

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.892140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.637393Z digest=sha256:50fadfe11037d66722580530ee5858df1040746ae163283381f483ad831d43b6

Observation 262e6737-2251-4969-8689-dc1ad8c3da03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.644209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.644209Z digest=sha256:2d2d09180e7cd7e1fe7f5a7c801899fea5af797500c07f7e3720332b3a92606d

Observation e5bd2610-ebc0-4921-beae-035751c68f56 · outbound

This paper cites Gaia2: Benchmarking.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Gaia2: Benchmarking

Reference 5

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.824200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.650845Z digest=sha256:bdbc4ae172b65266c8e5fbbd7ea224c8deacde0e0afea546240d98f50d3bd07b

Observation 5f2f6079-8e0c-423e-9418-3d4d740cc9cf · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.656643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.656643Z digest=sha256:c10d618b71341c122086859f25043a51180bd5e6b7523ab1d48b7a52e86177c9

Observation 2eed80e8-4fda-4798-b393-f5c232c1374f · outbound

This paper cites AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.662759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.662759Z digest=sha256:3b0426c86831afdeef2b29ccde95b4094708dc4a2d1a5e35f9bbcbb2913b627c

Observation aab3ff45-1e69-4b58-8a2b-06ba5ec7c50b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.671216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.671216Z digest=sha256:0fb86694832f666813ba4fd00ea1e514eb8a2f436cd86d182e6718ca419f87cb

Observation f6fa389a-5855-47e3-91a9-d72036d08e70 · outbound

This paper cites Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.642993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.676550Z digest=sha256:d823fe222bf54c151e1279dfba8148d15e844ff1849c6821a9c0bdbcc162d88c

Observation e2c12d1e-c277-49ae-9edd-0c9fd3d326f0 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.682339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.682339Z digest=sha256:c450f812f1c0dee5c7d70031ebe648c470c1eef3b90f1a109d68d643117c1343

Observation 1690fc54-620e-43d7-8ebc-5973c210db23 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.687685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.687685Z digest=sha256:308bd615689243c0c081e9b24066ec9a11ac82ad95db3bf2e4bb65747ca0e140

Observation e7fbe459-18b3-4e79-85f5-3ccb31a31ff5 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AppAgent: Multimodal Agents as Smartphone Users , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.693441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.693441Z digest=sha256:9527eaf8d201c0d0cc6a06e7a13145ada97f71fcc9893c86cff3e1ec3cb7ecd2

Observation 76666871-2092-490d-aa4e-e780c00513fe · outbound

This paper cites DeepSeek-V3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.698799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.698799Z digest=sha256:23aa564c504de7be8831ef6d3d3ff9724bbf0eaf8ec2e5198ad9ae12c8102cd4

Observation d908d904-43c2-43cd-8707-8af95370b703 · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.704359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.704359Z digest=sha256:eed81bb7aad752eb1bf01645a03228e146c95417de6e4b4a1c051441aac8a75b

Observation 3a3361d7-5ae0-455f-9895-ca98c1e57cfa · outbound

This paper cites AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.709584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.709584Z digest=sha256:61d9cee3a98174262e6d99b16a92586079a5d46c0c4215d544947b3825277447

Observation 3dda646a-c5d9-4723-85f3-7de59453ace0 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.715967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.715967Z digest=sha256:30a370e3918c7cb0a44b6b1f7fe28675ed6f85551a50e05f3a63edf926d79db7

Observation a46d83d4-d7ba-4869-ae16-463c5eba8be4 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.721936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.721936Z digest=sha256:196452f8581ec8dd242f4f7217ab0325ff6422cf0164bde46cb8c41cfcfa8aed

Observation 8999ef49-cbed-404e-9226-2314b38e0b14 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Twelfth International Conference on Learning Representations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.727753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.727753Z digest=sha256:48e4b2819f741087cab1f64cf6fe33929393730462002c8187b2cec1aedd69db

Observation 376c67ef-44bd-4260-9b91-fcdda83d8c95 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.733603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.733603Z digest=sha256:833640395f45364079895d434759412c66e362f8c7502efdc20dbd26ca23f0b5

Observation ab0b0730-1558-44c2-9e1e-c3ae8030677a · outbound

This paper cites FanOutQA:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FanOutQA:

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.739412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.739412Z digest=sha256:14b3ff3de40789e0ca7aefff1dafb588d001a800b45eef114b458967e68f17a8

Observation 98778e32-1a63-4115-b78c-204c22101d1d · outbound

This paper cites Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.746854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.746854Z digest=sha256:b704e6ab2d6d9ba7afd53cd20468b48e1d6458e677a2763e97d2119a066628ec

Observation ef4c0d86-c4b4-4050-bc2c-3b701dc087d9 · outbound

This paper cites Cohen and Ruslan Salakhutdinov and Christopher D.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Cohen and Ruslan Salakhutdinov and Christopher D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.753970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.753970Z digest=sha256:2fba808a1b6d3842b55539fcab4c203a3ea415d62dc634dfe961c92279b3d024

Observation 713569f4-462e-41df-8ac5-7ab811aef820 · outbound

This paper cites Generalizing Verifiable Instruction Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Generalizing Verifiable Instruction Following

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.763030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.763030Z digest=sha256:daddbc0deac26540be7f94230dce593c3855f7b57eef14fb6e5204a6c0f77d02

Observation f677e6b5-217f-4c6d-aa0a-463420e36a22 · outbound

This paper cites On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.770692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.770692Z digest=sha256:ecc7d7885993bb983d378bf10362845ff60764bd37d6a6014a19e9808fb0c63a

Observation 17883964-782b-490b-a663-6947cb8de8d1 · outbound

This paper cites Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =

Reference 25

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.778032Z digest=sha256:02ecdc68f78515d7a674fe58e6f7f3015db2e3de3449531947f601a971c0dd11

Observation 9b496bff-8d9a-478d-8414-54874be83004 · outbound

This paper cites TL-Training:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TL-Training:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.783860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.783860Z digest=sha256:3eeedeb9452e0af18e143c13162dfc3b55cc3187f29e9af4fb280423a72191cb

Observation a707554c-dccf-4d71-8ff1-4964b98cfabc · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Instruction-Following Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.790762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.790762Z digest=sha256:c3e51a153133b0477dc8c4c8fc4f410fb3ee0b30c35c5be3dd0269ba7d5769f6

Observation aacb158e-0f82-445c-a4aa-d300ef1634b9 · outbound

This paper cites First-Person Fairness in Chatbots.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information First-Person Fairness in Chatbots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.796199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.796199Z digest=sha256:82f3d29e7dde1931be2cd06b3b16a37f52d5c116f7a27b69ad83b52a54f219c2

Observation 504da331-f6d2-4bc9-8721-fc5ce57263aa · outbound

This paper cites Tool learning with large language models: a survey , journal =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Tool learning with large language models: a survey , journal =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.802775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.802775Z digest=sha256:3a61f3e3977d9884f1d2cd3d48955ba337488ebadf2959bde6578548bdc394db

Observation e8a15dec-dc31-4a63-9077-949d426406d9 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.808303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.808303Z digest=sha256:68910c75421409e391981ba69e52f0844d95724c8b7666d1b1556688ee9286f5

Observation 60fe019a-b7e9-434c-9d3c-66d7cf60f3aa · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 31

Resolution
parse uncertain
no resolver link, observed 2026-08-12T19:21:58.814046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.814046Z digest=sha256:c30502770c15f8c8e87a85c5b1908416c4638f2ab4b1f935ae3f19567d0d0f12

Observation 870b9e98-03eb-4190-9e6f-4636f8289b2e · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.819426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.819426Z digest=sha256:13fbbaf45d29b4e67d6cdcd4c0d6dcaefee32b8aa6799a682dd16c1856b8d11e

Observation 0f664a46-f911-4e3c-8dec-6054188e3f03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.825117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.825117Z digest=sha256:e0baf2888efcdd1936fc272d4c733e7ba2d38be43fe882f409c1face17ef6c4a

Observation 28bd6c66-ec43-40bb-a772-66eaeebf9336 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.830362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.830362Z digest=sha256:c1bbdc2463b010db4bf3395313f99f2cec402e89a9efad4bb2e09391d48d35e8

Observation 84bf6d1d-6af2-4e55-bd46-320b798ddfd3 · outbound

This paper cites Newell and P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Newell and P

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.836583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.836583Z digest=sha256:d5d6fa7c992cbb1c6d49c899430fa6bb2c559931811770994ac6544fe79c5c11

Observation cc89091e-2067-42c9-906d-0a82fbe3de2c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.842032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.842032Z digest=sha256:29b90d3269e71d9dabc07d7d717bb8798ef54d432250d8e4a80f91c1a6184cc4

Observation a5ecb3c6-38d4-400a-9903-416c564c4d3a · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.850170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.850170Z digest=sha256:8f2d92c53b274e034b71a6b786e16e2df16046c22c946156a2d6e667e101d30c

Observation 6e94bcfa-ef29-4123-8fa6-6c6146d3122b · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information and Stoica, Ion and Xing, Eric P

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.855533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.855533Z digest=sha256:028a12900cd048c676e768fa6935926bf5135d28899b03d340593c91283b8a59

Observation eead0c0a-dcdf-469c-a06c-a3382e82d138 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.862026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.862026Z digest=sha256:2246597eea93468d1edcddc6d5bf9e3344999e3f1ce9e4d123ed4dfdb1d6bd68

Observation f503f6ea-2063-4c3e-bdf7-4b9d802ef501 · outbound

This paper cites Scaling Laws for Neural Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Laws for Neural Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.868620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.868620Z digest=sha256:d79696ef5250d044a307c8551767493a92ca5da254079a3336fe5f3615f0543c

Observation 409a4d5f-eb31-4393-b95c-0ab7fa7dbbe1 · outbound

This paper cites 2011 , publisher=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2011 , publisher=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.875174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.875174Z digest=sha256:23072a82eb3475119f9f7546d91a93acc7d83e20a2a87a50b2573d00c7c51292

Observation d1bf6ce0-cf16-42e1-801e-9a1d8e4da587 · outbound

This paper cites RestGPT: Connecting Large Language Models with Real-World RESTful APIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RestGPT: Connecting Large Language Models with Real-World RESTful APIs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.880808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.880808Z digest=sha256:62ef77bc4146ffc5d30e2b53126e4af312d8b0e15c13923b98d5305df769b58b

Observation 677d6db5-626c-49a2-8c2c-1758ecae4f89 · outbound

This paper cites Improving Language Models by Retrieving from Trillions of Tokens , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Improving Language Models by Retrieving from Trillions of Tokens , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.886530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.886530Z digest=sha256:b769e75a9ec48e4cdd811148657d70037b904eadbaa74900516826bbb3fdc5b6

Observation 22ec1c83-0ad8-411f-8276-1f1df8abbe85 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.891620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.891620Z digest=sha256:546ce268f0a1febbd0c6426e72f59d95713ada9a5121ff587c4785553c7bdc25

Observation 559368ef-7b8c-4ab8-a315-6d5ec915e2f3 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.897223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.897223Z digest=sha256:88ed30c9d824cf2cfb85e49c76fff4dfa0c7c3bd0d89b9b9bdd3c3ea9c5b7358

Observation 76062a80-5d32-46be-b837-e9e9ab1b69fa · outbound

This paper cites TALM: Tool Augmented Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TALM: Tool Augmented Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.903319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.903319Z digest=sha256:3e6984d40154e55d9ad29d5101ccd917985c62ec4120c0a5f1c06b30f085db2a

Observation 16d20607-24fc-42bd-beca-d3ee04d9c111 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.908726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.908726Z digest=sha256:445309ade2e1260bf7573f3cd7ea78c13afbd8740d36335db0e9f7f99c79ffa4

Observation 6416dcc2-8d0e-40de-be70-bd8c56b2e984 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.914291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.914291Z digest=sha256:6f07dceced33b4c2aa80069689480503af1cb72f299140a93d8f70d8c4c204e9

Observation 0d8f8a41-0d7f-467f-9700-deacd409d182 · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.919706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.919706Z digest=sha256:6152103c8372dcd503c71cc1ea28f63d5e501d747f8829d1d00a007208935024

Observation 243aed30-d86c-4f25-98f4-3ca26e64c1db · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.925867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.925867Z digest=sha256:c2b0367b7dde899847c8dee11dee9906cf1772a05ec1ef681f1481417a956dbd

Observation 02c142c0-e3f2-4826-9212-702502787eff · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.932398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.932398Z digest=sha256:dc511096de85026a11a03e2c214eced3d99a92129d07e150c4a77f9b74d61a07

Observation 6a088319-384f-4d81-8edc-c2a88aa94bc5 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information WebGPT: Browser-assisted question-answering with human feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.938065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.938065Z digest=sha256:41a5c64e92494e839714d64baa0eaf605e190e3b8c4b48517b3e2c8e4db7937a

Observation 92a932cd-81a9-49f3-abcb-5181cffe8eaf · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Instruction-Finetuned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.944883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.944883Z digest=sha256:45a8ea5af54fc222b601366d74d9860652d644781f08dab4d6b42ff39bc42f99

Observation d0e7730e-4561-43e5-b1b9-12d4d2d64d5c · outbound

This paper cites Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.950479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.950479Z digest=sha256:96541b5e7175c7d8fc422c9afaf67f67d4a2ba37551dc24add6da330ec941876

Observation ffc3ef02-0713-4f13-9505-a6ef09f776e8 · outbound

This paper cites Narasimhan and Yuan Cao , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Narasimhan and Yuan Cao , title =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.956016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.956016Z digest=sha256:084330745068a600ce58f6ce0d975d4ffa44b8aa5494a6a33902f8e6feaf924e

Observation b290845d-1959-4e9d-aee2-e1cef7ef86fd · outbound

This paper cites Chi and Quoc V.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Quoc V

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.961736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.961736Z digest=sha256:f29cac25b5e885be0f6b3a9876cfb5720b963e643ab17fb06338bf702022ccf9

Observation 501493c6-4cda-46d7-af8f-2689efd2876d · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.967514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.967514Z digest=sha256:0937ad955a625e589a7b2a53460fd3cc0b02ad2fc5d0b94c1c101553e11baeae

Observation 18b28abb-e018-486c-b4ec-84c4705b8b48 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.973424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.973424Z digest=sha256:d72ad9afa036776a14ebde803f31be093f63074f9c7568ca4384f21246b486b2

Observation 43770c68-76d8-4d26-97d8-0aa098cde640 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.978942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.978942Z digest=sha256:66ae4e5885bb5832995b2a4453c3fce47187c12aff1d30432e1f8e427c1a15d3

Observation 67a41b55-0915-4817-b76f-f48acc6c4949 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.984416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.984416Z digest=sha256:c39b21b705cd5843f3b956b47e41e2ab1d49d7b592a1fabd0ca1b8814d604607

Observation aca258d9-31f2-48d8-8c41-62638fd2de37 · outbound

This paper cites , author=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information , author=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.990406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.990406Z digest=sha256:7e94e9fd8c4422f361d299e1d56918bfa000ccb0421dbfe1cd3c3845c6061dfd

Observation cd50d75b-b662-4e67-a8d2-d9eb00453e05 · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.996026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.996026Z digest=sha256:1177fc2bedeb2d1e0e05d881523f9bc80a8218ef7bc68d898af36163c26d1c80

Observation 444ba3f7-2efe-410e-9ed9-ad98eb8d3466 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.001037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.001037Z digest=sha256:24f43f050c1841f9748764086c7f800fd2e8dcb62e2af0b364a6b99af80670a2

Observation d61a495b-167e-4bd3-9e9c-9f728b833094 · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.008901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.008901Z digest=sha256:3a11d0951689f749ba87fd7bd1e7f447ae24632859be343d6ede05844226f22c

Observation 06fcce3a-1348-4f1a-8bbc-40eb06386d6c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.014624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.014624Z digest=sha256:29e563b97ef8b466dae82460303d8125e8bd58b964c8ff788b2dcf76518dbf60

Observation 7d6ee9cc-40a1-4fe8-ba2c-7d3bfa9bf6e4 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.019788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.019788Z digest=sha256:f7608a0f9150b0cc966a553a55f777546beb937fd9f55829a08da7bb89cabacc

Observation 7a84ff0a-1049-410a-ade2-7db6586bc985 · outbound

This paper cites Program Synthesis with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Program Synthesis with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.028908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.028908Z digest=sha256:ea8aa3350fb92a57c3e60ca233309de7ad8f44fb8b80a2153d04c39e288fac4f

Observation 5e6bd6cc-eee7-45d3-8a77-6595fb417be9 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Measuring Mathematical Problem Solving With the

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.034957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.034957Z digest=sha256:190aa1f4d148f87659da43a1bcf994d0636969b46cb4996255a87074bfebf22d

Observation c5f4779d-be23-4e95-a7ed-c12dbb3cf7a2 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.043085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.043085Z digest=sha256:88686f1ac6c4718abd0fc911c009ec39cfce1c445dc57ea5d245f1f872aa8c22

Observation 64e51fcf-4af9-40e3-abb9-65f4ea18934e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.049354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.049354Z digest=sha256:af480b85ea8121d0413eaf10d8a8ba24772c851f83a216efe153658103ce3486

Observation 3fec37c2-3645-4121-acad-f36ed77b903b · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.055383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.055383Z digest=sha256:6de43a5b0ab0f8fe7a04deee8dc4a60d42755b7e69c50f649440a7243ee23b76

Observation 43eecf78-d042-48cf-93b1-2fc5599aec3b · outbound

This paper cites IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.062541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.062541Z digest=sha256:d4c981db5152d15f0449ac6ccae686cb60fb3ff67f587c661212f3febaf8b1a0

Observation c381b3c8-f251-46ff-8408-0aeceb4b2af3 · outbound

This paper cites CIF-Bench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CIF-Bench:

Reference 73

Resolution
verified exact
doi, observed 2026-08-12T19:22:01.708390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T19:21:59.069885Z digest=sha256:9e238c05b93d664bcd63b7da10f5cdae23c5ee6a891a367d890fe23466b0c87d

Observation 219506e4-e879-48af-b797-41ff5f411b71 · outbound

This paper cites Manning and Stefano Ermon and Chelsea Finn , editor =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Manning and Stefano Ermon and Chelsea Finn , editor =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.077006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.077006Z digest=sha256:66da2f03b43252b9abb6726e81a38d21814ceffd4255ac890889843a8273b642

Observation e1fda741-9ab2-465b-8824-813dcb2ae246 · outbound

This paper cites FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.082243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.082243Z digest=sha256:f3e93b59a40d7b5829333c120af98bcc94014ee8d9e7b55b1902d602e879dd4e

Observation d7131c1c-784a-45c8-924e-ac75120185b5 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.087672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.087672Z digest=sha256:4ee060ae37e07a91018479fed8e1e93eedbefccc89dc92b169c0a61fe40231f6

Observation 37c7d433-dc3e-4ac7-a205-acaa6e91b16b · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.093215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.093215Z digest=sha256:7b2bda9b47023c7489776d2b060034adc46dc38dda355a6d277602b4c9497b4d

Observation 89c0deb6-04e4-46e9-a823-fd598e39e6a4 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.098498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.098498Z digest=sha256:6a115b3ca4e62acb2cf755df18039ae5a8d83d6d43cffe4de99a01420f513c10

Observation bb77a871-757a-4ad2-842f-b800da01771a · outbound

This paper cites SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.104812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.104812Z digest=sha256:39119760db1bd6fd9955a2b654f905be27787934ad414804ef5d310b04b6cd69

Observation 7cb0ccac-6e5e-4765-8b12-d6536c089f2e · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.112477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.112477Z digest=sha256:427e2a983db15fb09fd655818dcc43284d51117cb5d3cea22681dc0a113f1881

Observation b21da20d-4005-4b86-bb06-3eee258a38e9 · outbound

This paper cites FollowBench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowBench:

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.118087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.118087Z digest=sha256:8a06fb84723b294a6c2c70534e996877609107c5b83c8e305993f8d6ecbffe3e

Observation 8dea1378-933a-4d8f-856b-792e97303e34 · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Needle in the Haystack for Memory Based Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.129521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.129521Z digest=sha256:edc6f4a3eca1014e53d91e4314d6b074f0868142ee491c181fa12910a689c21c

Observation f76cd0a4-0968-40aa-a153-c79214cd5fc1 · outbound

This paper cites Qwen3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen3 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.135717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.135717Z digest=sha256:f477013d50882f9eb5ebe2dafbc11c156ed7481b8b2650bbef2f04e2693a0963

Observation 259dd41d-7df7-40b7-9726-859a10de69a5 · outbound

This paper cites Findings of the Association for Computational Linguistics,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Findings of the Association for Computational Linguistics,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.143832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.143832Z digest=sha256:00bd3abad52fc77d4bd3642b31db612655bcc96e37d6c3c8d2a16b9ca9710ec6

Observation dc5eecb0-ae1b-4ca8-b5ad-b73c7525b84c · outbound

This paper cites InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.152621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.152621Z digest=sha256:a7464a59186a3240c46eddb89fd98caa8a1b4d979d5eecd0304133d53e4e17ed

Observation e6d96620-9d51-4b89-af16-2dd3598b649a · outbound

This paper cites Long Context vs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Long Context vs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.160015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.160015Z digest=sha256:0a243b608eb96a1b7c88b99f67ef3f8cf6fc9f8c721b781329ca44db5be0e7d8

Observation c183f780-9786-431c-b428-3cd84bbbf6d9 · outbound

This paper cites MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.166468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.166468Z digest=sha256:bdd01270c93ffbc06ea75373186fe52e24d261565681e85aed09bc0fd4c5437e

Observation f43e3356-ddf2-45e9-889f-ebe1c070661a · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.172188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.172188Z digest=sha256:4c5737dfd22bc1fe3840da30f5f5cf7b508bb1053d52197f13e1e78d9c8965fe

Observation 6edc0562-3a95-46b9-8612-4b61d88ba15b · outbound

This paper cites Hanjie and Runzhe Yang and Karthik R.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Hanjie and Runzhe Yang and Karthik R

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.179579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.179579Z digest=sha256:b67ded3a3aa96bc9963fc387847e779f1aff64c5dc8cd3a98aef8fefed5d4fe0

Observation adc69345-6af4-49d4-be9c-d20a6219a7af · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-OCR: Contexts Optical Compression

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.189384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.189384Z digest=sha256:34f3fa040129bad8f45ff23ce35b08caa39794824a91adc7ccba0f80eb5a828f

Observation 0a24fd01-5bb2-49dc-9b7f-205d925dea79 · outbound

This paper cites Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.195991Z digest=sha256:910efc66b95a36f20901d5598733bf15f5f8a7e7d9738dafd3af1b61bdd1e555

Observation 2c123124-ae6f-400b-ad14-e6cb13f5a139 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.203762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.203762Z digest=sha256:f4786df2ec35d0ca685c1ec8f20c310ba729ae8a1b4d545b93ea6b352ad84054

Observation 10870325-622c-473c-b126-934be6591ef0 · outbound

This paper cites Qwen2.5 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen2.5 Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.213050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.213050Z digest=sha256:3e63b3d2ace987b712687c082d26f39fad9310d4d9483576aa765c1331e95f32

Observation b581982a-1e51-4025-b3d5-e08152b243f8 · outbound

This paper cites 2024 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2024 , url=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.219235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.219235Z digest=sha256:987b75472e5cd9af4014db48fb19b51f59dd367b37827ea257e386ff16e6c76d

Observation 478158d4-3e28-4940-9f30-97e5f055682d · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.225271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.225271Z digest=sha256:3891389f19d9e7f6db8839bbe8138885f6f001de6504b0f14158896212abb1d3

Observation 24e30cd2-de56-4bd5-8032-d86f246dcfc0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Code Llama: Open Foundation Models for Code

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.232907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.232907Z digest=sha256:2d001720ae31201bb7ee2ccdd9200b0b71152a00cf98d970b8421c86590c8772

Observation 3a52fc40-6cf4-4664-a5e6-d87f60e13b2b · outbound

This paper cites ToolHop:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolHop:

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.238846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.238846Z digest=sha256:def1f05687819104b9a660a7bdc67046f9a44612886850979f1761309aa8947f

Observation bd455ddb-5ec9-46e5-8f49-1a32cc7c7b27 · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.245764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.245764Z digest=sha256:f7c68f45b8fac84d2c40a2f348c4b94e9e86ad72cd80e18ba60fabc0c4331b7b

Observation 44819e06-aebc-406c-9097-bb6d2c07442b · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.252278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.252278Z digest=sha256:df8ec6f06c4b364a46c2cb3846f7e3e5429c48fe33749586aeb690d930da7254

Observation 4081df13-c0af-4874-95c5-8689dd87b6c1 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.257857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.257857Z digest=sha256:fe7e3bcc5e6f7b8a51337bb40fd7093fa46ab4e061020dca22578986f25a2da7

Pith citing papers

No inbound Pith citation observations are available.