Pith. sign in

Paper Citation Record · LEDGER

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

As of 21 August 2026, this Paper Citation Record lists 100 of 218 outbound references and 0 inbound Pith citation observations for arXiv:2608.10692.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10692 v1

Coverage vector

measured 100 of 218 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.257857Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 218 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved94
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d3d2b53-cf36-4553-99ae-37d9bec5e39a · outbound

This paper cites Langley , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Langley , title =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.624053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.624053Z digest=sha256:eda845410378fba20c6ce79c7936bfeb94c42fdadf9f3e4c33dd4cae71601977

Observation 2e475699-9ac0-4f50-ab64-29a8dff91f87 · outbound

This paper cites AutoGLM: Autonomous Foundation Agents for GUIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AutoGLM: Autonomous Foundation Agents for GUIs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.630665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.630665Z digest=sha256:353a72c79e8add68093212790e0ed66e788d94ef8ef5bba0854ac5336eb23a6c

Observation 2cbf283e-58e2-4f43-a012-a5ff84ee17f4 · outbound

This paper cites Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis , booktitle =

Reference 3

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.892140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.637393Z digest=sha256:5d369aa97e1b551fbfc9b7481cba2f512a62e688530f4cc4b5f64eac2518e4c4

Observation 262e6737-2251-4969-8689-dc1ad8c3da03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.644209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.644209Z digest=sha256:79d7f4ca001a0e9e124f4a4fc839abb266decf9aa16b357be10c6b5004372a46

Observation e5bd2610-ebc0-4921-beae-035751c68f56 · outbound

This paper cites Gaia2: Benchmarking.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Gaia2: Benchmarking

Reference 5

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.824200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.650845Z digest=sha256:e0cc483b013998f222ac21d09ce654d3ed163c3b754be86e89327bd6924738ba

Observation 5f2f6079-8e0c-423e-9418-3d4d740cc9cf · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.656643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.656643Z digest=sha256:b08069adf93c2f2545e06cdee9b8fdb7dab9310c27120616ca1e04f09dccbb6d

Observation 2eed80e8-4fda-4798-b393-f5c232c1374f · outbound

This paper cites AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents , booktitle =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.662759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.662759Z digest=sha256:16716323210de30447c29cf159bbd19682c8ef0d156db3ad87e6d98b6dd21498

Observation aab3ff45-1e69-4b58-8a2b-06ba5ec7c50b · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.671216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.671216Z digest=sha256:ceafcacb982df1d010a06d7e620506cb0f3ca4dc876a6f5d40b5ea1f5bb22fd8

Observation f6fa389a-5855-47e3-91a9-d72036d08e70 · outbound

This paper cites Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents , booktitle =

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.642993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.676550Z digest=sha256:2fb7a169bd48e18489268ef16bb76af9d853d577b54429c0acaabd9fc642761e

Observation e2c12d1e-c277-49ae-9edd-0c9fd3d326f0 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.682339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.682339Z digest=sha256:c033a4abcdda6cc30c691be4a10133c6f4a818c0fe96f2a3dcf4a59daf8490b5

Observation 1690fc54-620e-43d7-8ebc-5973c210db23 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.687685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.687685Z digest=sha256:722975bee19f03838ea0b1fd3d7e9943f62b9953f19e487e6d6b88e705a6c48f

Observation e7fbe459-18b3-4e79-85f5-3ccb31a31ff5 · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AppAgent: Multimodal Agents as Smartphone Users , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.693441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.693441Z digest=sha256:2b1c561db8ae5dd761ee12ecdf1539c56fe49b4d0cec55cf0207d04491c7020b

Observation 76666871-2092-490d-aa4e-e780c00513fe · outbound

This paper cites DeepSeek-V3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.698799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.698799Z digest=sha256:048fb2f7e5e9aec12eb0cdd72d9933b90fcaa33b26a130bda4fddcafcdfc5213

Observation d908d904-43c2-43cd-8707-8af95370b703 · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.704359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.704359Z digest=sha256:e8c1b915e94eecf4c7a67acb072abf46cc042fdcde2932606a92fe70d5101d01

Observation 3a3361d7-5ae0-455f-9895-ca98c1e57cfa · outbound

This paper cites AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.709584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.709584Z digest=sha256:d364fb6503b64ab7b6bda424ddfe6a4c3bd8f3d30d5c2bbe3e1e11b6aee20773

Observation 3dda646a-c5d9-4723-85f3-7de59453ace0 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.715967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.715967Z digest=sha256:a76e44593feb7d022781bd77a6965b824d706d46e2fae1c6a0cfa0fbfa5a15bc

Observation a46d83d4-d7ba-4869-ae16-463c5eba8be4 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.721936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.721936Z digest=sha256:9a904e5cf4ec14d31c3a9bb71a824d13deed9194f885f36d197f9c15bbcb5ed2

Observation 8999ef49-cbed-404e-9226-2314b38e0b14 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Twelfth International Conference on Learning Representations,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.727753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.727753Z digest=sha256:951c66ef4db62e8b149f5e95f3a529a20a1c9d2960ce2fcd0f5dea88018bb2e7

Observation 376c67ef-44bd-4260-9b91-fcdda83d8c95 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.733603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.733603Z digest=sha256:2b13bbed7038414e042b838a364714bff3676f667c6d4bef8e7d6a28640df647

Observation ab0b0730-1558-44c2-9e1e-c3ae8030677a · outbound

This paper cites FanOutQA:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FanOutQA:

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.739412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.739412Z digest=sha256:7e4cb890793691d29e0fc00a79115050f21e3d180a90ce705ce982f0d0c6463c

Observation 98778e32-1a63-4115-b78c-204c22101d1d · outbound

This paper cites Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Evaluating Design Decisions for Dual Encoder-based Entity Disambiguation , booktitle =

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.746854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.746854Z digest=sha256:bf10e1126aba007ce7c821325ec1205f028836c067fc1a4e0e9c4a19d4b75ec1

Observation ef4c0d86-c4b4-4050-bc2c-3b701dc087d9 · outbound

This paper cites Cohen and Ruslan Salakhutdinov and Christopher D.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Cohen and Ruslan Salakhutdinov and Christopher D

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.753970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.753970Z digest=sha256:3b2e8fd1e57d2b45317f81a09d0fbc5da5f6c9c76602a4b0f19ed0971ecb8501

Observation 713569f4-462e-41df-8ac5-7ab811aef820 · outbound

This paper cites Generalizing Verifiable Instruction Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Generalizing Verifiable Instruction Following

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.763030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.763030Z digest=sha256:d5d239ea1493e07be27a50e0148bc8053d8cfb42dcbd4b1197746f45486522f2

Observation f677e6b5-217f-4c6d-aa0a-463420e36a22 · outbound

This paper cites On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Multi-turn Instruction Following for Conversational Web Agents , booktitle =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.770692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.770692Z digest=sha256:51479a464fd1d9cc27f8affcd272fbf364a804db99abafde8d8683d50ad5789f

Observation 17883964-782b-490b-a663-6947cb8de8d1 · outbound

This paper cites Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toward Verifiable Instruction-Following Alignment for Retrieval Augmented Generation , booktitle =

Reference 25

Resolution
verified exact
doi, observed 2026-08-12T19:23:28.426948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:21:58.778032Z digest=sha256:2ac5f957713352f2ee0acf4b43d95baafd775197ffc84ded3a1de4a08f39a919

Observation 9b496bff-8d9a-478d-8414-54874be83004 · outbound

This paper cites TL-Training:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TL-Training:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.783860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.783860Z digest=sha256:624563acaa43f157c656d845065b57a5826e7b1b43436fbbb074c28608b6585b

Observation a707554c-dccf-4d71-8ff1-4964b98cfabc · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Instruction-Following Evaluation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.790762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.790762Z digest=sha256:2b0c33303b75042964953569b35b97623f55dcff11c0ab9676a2211334d030b0

Observation aacb158e-0f82-445c-a4aa-d300ef1634b9 · outbound

This paper cites First-Person Fairness in Chatbots.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information First-Person Fairness in Chatbots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.796199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.796199Z digest=sha256:41e5fda27774bfe48221f5b20f1ddde6a7ab963dc850b784ca8da1cdfb623a2b

Observation 504da331-f6d2-4bc9-8721-fc5ce57263aa · outbound

This paper cites Tool learning with large language models: a survey , journal =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Tool learning with large language models: a survey , journal =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.802775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.802775Z digest=sha256:784076a80fd509c030776fd55bcc5efc04119c68d195abb564c13d68faa42f5a

Observation e8a15dec-dc31-4a63-9077-949d426406d9 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.808303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.808303Z digest=sha256:fa9ecbe3eec364032b0e07795f31c1e615f0de002ddd4e92f1c7cc193652d55d

Observation 60fe019a-b7e9-434c-9d3c-66d7cf60f3aa · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 31

Resolution
parse uncertain
no resolver link, observed 2026-08-12T19:21:58.814046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.814046Z digest=sha256:0cc658c5434158fed69e7d02615e8c22745ea4a38f49c54570deedf9cb2cc4f7

Observation 870b9e98-03eb-4190-9e6f-4636f8289b2e · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.819426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.819426Z digest=sha256:0ff57268c78637d96a09b6ad6e737b33fe0154902e16ce8b04b203c33d97dadd

Observation 0f664a46-f911-4e3c-8dec-6054188e3f03 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.825117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.825117Z digest=sha256:6baa236fe3479b1547691b035bd07fddbc88649d4bebec6ffca8c59a2e78c736

Observation 28bd6c66-ec43-40bb-a772-66eaeebf9336 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.830362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.830362Z digest=sha256:58feb6d577978d009f0cf3e55c671c9ff7a4fe617507dc8d3cbcf830b5baef2a

Observation 84bf6d1d-6af2-4e55-bd46-320b798ddfd3 · outbound

This paper cites Newell and P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Newell and P

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.836583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.836583Z digest=sha256:398bc5c36ba38b4410747f9d4868c2f84ebeb6ff769fe4b417a21781c243d41e

Observation cc89091e-2067-42c9-906d-0a82fbe3de2c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.842032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.842032Z digest=sha256:f463ca008889a2557cdc38f9128785d97547e32551b3ae84d03de88a952965e6

Observation a5ecb3c6-38d4-400a-9903-416c564c4d3a · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Toolformer: Language Models Can Teach Themselves to Use Tools , booktitle =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.850170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.850170Z digest=sha256:6442ccd74d42282463bc359702b9711ddef0c6db0d4a93c4aba8278e32aa31b6

Observation 6e94bcfa-ef29-4123-8fa6-6c6146d3122b · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information and Stoica, Ion and Xing, Eric P

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.855533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.855533Z digest=sha256:5ec80c4f6dc1a608f5492e5417de66a9df2984b81c6ffd3e8250a8a25ace58fc

Observation eead0c0a-dcdf-469c-a06c-a3382e82d138 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.862026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.862026Z digest=sha256:fcff5c4a8697dc25fa80f7929b1c66d18ff95121d27c9215c3ef5344011e99b6

Observation f503f6ea-2063-4c3e-bdf7-4b9d802ef501 · outbound

This paper cites Scaling Laws for Neural Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Laws for Neural Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.868620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.868620Z digest=sha256:9cd0707b2c02e3e93fca1fe943a5b94f30738306ea4248de8c6050e2ffec318d

Observation 409a4d5f-eb31-4393-b95c-0ab7fa7dbbe1 · outbound

This paper cites 2011 , publisher=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2011 , publisher=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.875174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.875174Z digest=sha256:3d400606c96425f094c0341f3fd5cbe3c4a789ec9c91051b57be12cc43fa00e3

Observation d1bf6ce0-cf16-42e1-801e-9a1d8e4da587 · outbound

This paper cites RestGPT: Connecting Large Language Models with Real-World RESTful APIs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RestGPT: Connecting Large Language Models with Real-World RESTful APIs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.880808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.880808Z digest=sha256:05c1148980defab60e32d53b4e5fd99aa867bc31d8024f8c6fa76c60c0ebb9b1

Observation 677d6db5-626c-49a2-8c2c-1758ecae4f89 · outbound

This paper cites Improving Language Models by Retrieving from Trillions of Tokens , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Improving Language Models by Retrieving from Trillions of Tokens , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.886530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.886530Z digest=sha256:48772dba24201e04b6c2673862274b72d28318a77a188c3c03975d0434e4211a

Observation 22ec1c83-0ad8-411f-8276-1f1df8abbe85 · outbound

This paper cites Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.891620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.891620Z digest=sha256:1982646a07266e90ea18b5c44ae8a5f655631053581bf925965848d18581da07

Observation 559368ef-7b8c-4ab8-a315-6d5ec915e2f3 · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information LaMDA: Language Models for Dialog Applications

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.897223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.897223Z digest=sha256:f55143fd527ca33faa965ae59e102647013884a94f459ed99d945485fd93d56c

Observation 76062a80-5d32-46be-b837-e9e9ab1b69fa · outbound

This paper cites TALM: Tool Augmented Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TALM: Tool Augmented Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.903319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.903319Z digest=sha256:58c99eb890d99ae867a125074775850777de87378ba56fe8299bf4c5f90533e0

Observation 16d20607-24fc-42bd-beca-d3ee04d9c111 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.908726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.908726Z digest=sha256:7de475e8872cb0b22db516589ac87f29257da77e2787bd61873541bbce3bb276

Observation 6416dcc2-8d0e-40de-be70-bd8c56b2e984 · outbound

This paper cites ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.914291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.914291Z digest=sha256:96badfec44e019aa84bb71718b369e7435c77aa7b77358957ebe43fc1ddb4f1a

Observation 0d8f8a41-0d7f-467f-9700-deacd409d182 · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information On the Tool Manipulation Capability of Open-source Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.919706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.919706Z digest=sha256:29fe381ab550195a5a79d43a76eba89733320a981b488061dbab9da4eea7bc4b

Observation 243aed30-d86c-4f25-98f4-3ca26e64c1db · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.925867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.925867Z digest=sha256:da0fb324ac18250e67a5290cd77e5d7c952e06f60255e69b95eccf5994dd7262

Observation 02c142c0-e3f2-4826-9212-702502787eff · outbound

This paper cites CoRR , volume =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CoRR , volume =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.932398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.932398Z digest=sha256:b575721eaa1e2ad77cec223ecbdcbc676a65be60e7fd2e1b92b8968f44f4c74c

Observation 6a088319-384f-4d81-8edc-c2a88aa94bc5 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information WebGPT: Browser-assisted question-answering with human feedback

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.938065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.938065Z digest=sha256:1160f8482056d31ef029dbd38c627809e95185d049ef4abb438a4967b5de3dd1

Observation 92a932cd-81a9-49f3-abcb-5181cffe8eaf · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Scaling Instruction-Finetuned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.944883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.944883Z digest=sha256:fcfa22b9512e1d0ab799e1b2ba7ed6eed598d0fb37e7afda74c92acd2cf0ae44

Observation d0e7730e-4561-43e5-b1b9-12d4d2d64d5c · outbound

This paper cites Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Tatsunori Hashimoto and Oriol Vinyals and Percy Liang and Jeff Dean and William Fedus , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.950479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.950479Z digest=sha256:a919d83d42c47bd4cb8f5cf40525338354115f8762f6830df1cd4711bea9e697

Observation ffc3ef02-0713-4f13-9505-a6ef09f776e8 · outbound

This paper cites Narasimhan and Yuan Cao , title =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Narasimhan and Yuan Cao , title =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.956016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.956016Z digest=sha256:9c1bf1bed04421cf5609a316a859cff66eaff346a248f53bbfc67bebc1e28294

Observation b290845d-1959-4e9d-aee2-e1cef7ef86fd · outbound

This paper cites Chi and Quoc V.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Chi and Quoc V

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.961736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.961736Z digest=sha256:e8b1f92550bbbee6dbfd89aa2941e45e97f4fe67298d388046cf41592eacd4dd

Observation 501493c6-4cda-46d7-af8f-2689efd2876d · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.967514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.967514Z digest=sha256:d1292b9499274a86bdc57ecf9916fa55278ca775a1de0b3a9e130925b8f4d960

Observation 18b28abb-e018-486c-b4ec-84c4705b8b48 · outbound

This paper cites GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.973424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.973424Z digest=sha256:8a6354043731286da6b4adeb2c90e1d5c5ce91e9d0d27f72a827503d8d0a4fec

Observation 43770c68-76d8-4d26-97d8-0aa098cde640 · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.978942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.978942Z digest=sha256:345c38bd26abd4ac54b995519d2dbdd2633a6afcb47c47d776d06966079bf39b

Observation 67a41b55-0915-4817-b76f-f48acc6c4949 · outbound

This paper cites GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction , booktitle =

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.984416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.984416Z digest=sha256:3b8fdf06e63bb30435b67fb991f7812e13ed0f88a5168e08503c922b2874c9cb

Observation aca258d9-31f2-48d8-8c41-62638fd2de37 · outbound

This paper cites , author=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information , author=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.990406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.990406Z digest=sha256:ab2ca84cc8531db4b4dbd2d057e57535a10e6c00e31e8a2e990aed65b591a2ef

Observation cd50d75b-b662-4e67-a8d2-d9eb00453e05 · outbound

This paper cites Patil and Tianjun Zhang and Xin Wang and Joseph E.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Patil and Tianjun Zhang and Xin Wang and Joseph E

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:58.996026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:58.996026Z digest=sha256:b59f0ef29f94f9f3f9c8e29af05a85e0815ef5e3f036fe871daa3c80b82a5861

Observation 444ba3f7-2efe-410e-9ed9-ad98eb8d3466 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.001037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.001037Z digest=sha256:c666ed5fa742b40ef957ebc8d373035dbb4916ec8eb3b1cd0311f778dc4fea18

Observation d61a495b-167e-4bd3-9e9c-9f728b833094 · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning , booktitle =

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.008901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.008901Z digest=sha256:2f697a30852668d42eb3cb52df24ec4217e3157649fcdde6a8aac13762a626e3

Observation 06fcce3a-1348-4f1a-8bbc-40eb06386d6c · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.014624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.014624Z digest=sha256:545d5862f2b71c718c379039a9a90a357444d130fce55b4455b4254ad5043ff0

Observation 7d6ee9cc-40a1-4fe8-ba2c-7d3bfa9bf6e4 · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? , booktitle =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.019788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.019788Z digest=sha256:8cdd3ad5128412d7594a1300cd0c6a9cd815cd2a2277f6ac47ec0aefaea32368

Observation 7a84ff0a-1049-410a-ade2-7db6586bc985 · outbound

This paper cites Program Synthesis with Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Program Synthesis with Large Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.028908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.028908Z digest=sha256:b961f0836ec8622c9176e68322ccfc9c2ab1f804ad5a0448a97f10e44db4c418

Observation 5e6bd6cc-eee7-45d3-8a77-6595fb417be9 · outbound

This paper cites Measuring Mathematical Problem Solving With the.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Measuring Mathematical Problem Solving With the

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.034957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.034957Z digest=sha256:3024809d758d0f6c29618bf812af6cc35271c04ba30c8afe84e7467c0c79eb10

Observation c5f4779d-be23-4e95-a7ed-c12dbb3cf7a2 · outbound

This paper cites an unresolved cited work.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.043085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.043085Z digest=sha256:448c8c7e0228a6c925bbfba5c0ff55d3b3ba057261c456100b9cfa8dbacbe764

Observation 64e51fcf-4af9-40e3-abb9-65f4ea18934e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.049354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.049354Z digest=sha256:f0d261eef25651740956c76ca89100c68bf6b1b98521565c96c41ecd7372bc13

Observation 3fec37c2-3645-4121-acad-f36ed77b903b · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InFoBench: Evaluating Instruction Following Ability in Large Language Models , booktitle =

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.055383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.055383Z digest=sha256:cb4522a59e2688737a79a649450ed759988c23477335cf79c6b7c051997952ea

Observation 43eecf78-d042-48cf-93b1-2fc5599aec3b · outbound

This paper cites IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information IOPO: Empowering LLMs with Complex Instruction Following via Input-Output Preference Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.062541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.062541Z digest=sha256:719c2eee9b63af5a75cfc162aa45bf30f41931853b9b34141ea6f2f362118a51

Observation c381b3c8-f251-46ff-8408-0aeceb4b2af3 · outbound

This paper cites CIF-Bench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information CIF-Bench:

Reference 73

Resolution
verified exact
doi, observed 2026-08-12T19:22:01.708390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-12T19:21:59.069885Z digest=sha256:9be2e33bbf15be1fe0c984bc49569cc7da98ac2078bc5a684123b917937567ec

Observation 219506e4-e879-48af-b797-41ff5f411b71 · outbound

This paper cites Manning and Stefano Ermon and Chelsea Finn , editor =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Manning and Stefano Ermon and Chelsea Finn , editor =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.077006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.077006Z digest=sha256:4bdaa68ee7e791e04170e3804f7c3d258f1d6608b69c3c99ec41a75309fd17b9

Observation e1fda741-9ab2-465b-8824-813dcb2ae246 · outbound

This paper cites FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.082243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.082243Z digest=sha256:36b3e6cfb5886c4270a1132b032fdd0f2231a87e1d6de4fd184368e7c28d5f2c

Observation d7131c1c-784a-45c8-924e-ac75120185b5 · outbound

This paper cites From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models , booktitle =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.087672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.087672Z digest=sha256:ff33797fa40d5ec562b8a8dbf729fb2220d2a842c0804bbb81393e2dd18c514c

Observation 37c7d433-dc3e-4ac7-a205-acaa6e91b16b · outbound

This paper cites Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.093215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.093215Z digest=sha256:c9da686a30361dac4c101bccd93cb7a1d1b588836d62298503b8c5c6668194e4

Observation 89c0deb6-04e4-46e9-a823-fd598e39e6a4 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.098498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.098498Z digest=sha256:b647b76336d1aaca79c99403db029b596f4cb5d679beb4f1263437593f0ad88e

Observation bb77a871-757a-4ad2-842f-b800da01771a · outbound

This paper cites SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.104812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.104812Z digest=sha256:85b8709abd7f09b34cd2a188d7ed332d1fac96d6cd2adf981cd461cace172b84

Observation 7cb0ccac-6e5e-4765-8b12-d6536c089f2e · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.112477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.112477Z digest=sha256:c9732356ee7c8071ca50db089d14fd9ed551c9693b138188be578d4327f338c5

Observation b21da20d-4005-4b86-bb06-3eee258a38e9 · outbound

This paper cites FollowBench:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowBench:

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.118087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.118087Z digest=sha256:b282d9ba1b2e53a7b44261f2758710d28f19d7eb758d27c003e4331836d8f711

Observation 8dea1378-933a-4d8f-856b-792e97303e34 · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Needle in the Haystack for Memory Based Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.129521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.129521Z digest=sha256:020a7c64930f2f5095cad0cb5d2bc7620cc6f2d4047f7146dca616dcd188955c

Observation f76cd0a4-0968-40aa-a153-c79214cd5fc1 · outbound

This paper cites Qwen3 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen3 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.135717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.135717Z digest=sha256:04af8f04c6dfe54b959966211abb0adb6e65e867e4a6bffa4cc23471e140ecd4

Observation 259dd41d-7df7-40b7-9726-859a10de69a5 · outbound

This paper cites Findings of the Association for Computational Linguistics,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Findings of the Association for Computational Linguistics,

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.143832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.143832Z digest=sha256:23e1a6ec210cf7a026ec7ffb99ffabe00b6966b6117a068abb449ef8262a493f

Observation dc5eecb0-ae1b-4ca8-b5ad-b73c7525b84c · outbound

This paper cites InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information InstructUIE: Multi-task Instruction Tuning for Unified Information Extraction

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.152621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.152621Z digest=sha256:3f74c1b2169523eb80313f48a870423eeef88ebb3a167a6d632aee4a425c32c1

Observation e6d96620-9d51-4b89-af16-2dd3598b649a · outbound

This paper cites Long Context vs.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Long Context vs

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.160015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.160015Z digest=sha256:ee6e782c0ec605f8826e8ff06cc2377f8f18ed54020feb0f803509836376cfcc

Observation c183f780-9786-431c-b428-3cd84bbbf6d9 · outbound

This paper cites MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.166468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.166468Z digest=sha256:dd6c1ab418b345c27ac722f705bed4d8eb66f0a63b7b7bcaa99058d603a13d4b

Observation f43e3356-ddf2-45e9-889f-ebe1c070661a · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information The Thirteenth International Conference on Learning Representations,

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.172188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.172188Z digest=sha256:b8f42ed5dfbab9e972a96fb37b3945bfa5587eb93cd85901a653dbc5d9db5dba

Observation 6edc0562-3a95-46b9-8612-4b61d88ba15b · outbound

This paper cites Hanjie and Runzhe Yang and Karthik R.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Hanjie and Runzhe Yang and Karthik R

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.179579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.179579Z digest=sha256:a98308e1692395c45866df18d08866187fd29059e9911229fcf20c9f33523051

Observation adc69345-6af4-49d4-be9c-d20a6219a7af · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-OCR: Contexts Optical Compression

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.189384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.189384Z digest=sha256:8c28272f00a2ec127058ebfc17fce27091884e0ce6d72a2f8f6e8a9a59fc0d93

Observation 0a24fd01-5bb2-49dc-9b7f-205d925dea79 · outbound

This paper cites Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environments

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.195991Z digest=sha256:6ea7bcc5d320650ca6d8a4486c3a0cc408552bc99b2b5be75ec33c0906f38c63

Observation 2c123124-ae6f-400b-ad14-e6cb13f5a139 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.203762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.203762Z digest=sha256:b62f25b8b1cc78016727a23a49bd6ed3444197eaaef8ad39b3f98aba9188e185

Observation 10870325-622c-473c-b126-934be6591ef0 · outbound

This paper cites Qwen2.5 Technical Report.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Qwen2.5 Technical Report

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.213050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.213050Z digest=sha256:395257221502df0968f716a03b5ca1035aa9ad09e33533f71a7686bc35e9918d

Observation b581982a-1e51-4025-b3d5-e08152b243f8 · outbound

This paper cites 2024 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2024 , url=

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.219235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.219235Z digest=sha256:e95dc64adda6e92b1d385ef24af7647c32a2a29eb446990dd6b6787bfc7a7aca

Observation 478158d4-3e28-4940-9f30-97e5f055682d · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.225271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.225271Z digest=sha256:174c7989c5aa250c221b5cde8d7abb008056fb7378deea212e0cf8f9c796a6b4

Observation 24e30cd2-de56-4bd5-8032-d86f246dcfc0 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Code Llama: Open Foundation Models for Code

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.232907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.232907Z digest=sha256:35d39368f940a4e55eb5e708ade82cc2e8e04073fbfaa3f6d7778e2ba9b895c7

Observation 3a52fc40-6cf4-4664-a5e6-d87f60e13b2b · outbound

This paper cites ToolHop:.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information ToolHop:

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.238846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.238846Z digest=sha256:48d22e99138b75cd187c2be0572f488495e0d86e614b7ef73d796420b843ee70

Observation bd455ddb-5ec9-46e5-8f49-1a32cc7c7b27 · outbound

This paper cites 2023 , url=.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information 2023 , url=

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.245764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.245764Z digest=sha256:5ffa79139791d29979d50a570816b9beafb622f6cedd8a3c5aee879e4c73dce0

Observation 44819e06-aebc-406c-9097-bb6d2c07442b · outbound

This paper cites T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step , booktitle =

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.252278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.252278Z digest=sha256:8b916448ffad30d16599da1db68179ff2141d3c48f5c2878a0957afc4afb6807

Observation 4081df13-c0af-4874-95c5-8689dd87b6c1 · outbound

This paper cites TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information TaskBench: Benchmarking Large Language Models for Task Automation , booktitle =

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.257857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.257857Z digest=sha256:f87b05f9a3f6c633e6540502fbdf01c77e568d0541a77b0df1ab2b200af7e60e

Pith citing papers

No inbound Pith citation observations are available.