Pith. sign in

Paper Citation Record · LEDGER

Training a Vision Language Model as Smartphone Assistant

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 2 inbound Pith citation observations for arXiv:2404.08755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.08755 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 2 of 2 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:16:11.595838Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 697a33a3-ee73-4127-9afc-bd11d0624b55 · inbound

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions cites this paper.

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions Training a Vision Language Model as Smartphone Assistant

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.306200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T04:59:36.994758Z digest=sha256:e436b236a56c99271a93b8493a1de0f8a7cde923b3668e132827f66c4cdf4ce5

Observation 5fe55f0f-b8cb-4772-9820-5fa2c5d4dd13 · inbound

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data cites this paper.

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data Training a Vision Language Model as Smartphone Assistant

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:16:11.595838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:16:11.595838Z digest=sha256:949a39bbf70293f640a2f0994f47ed6916bee1e70c214cc77f1f253c63258575