Pith. sign in

Paper Citation Record · LEDGER

JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.16365.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.16365 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:40.483549Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5854a8e8-56d2-44d4-8836-6ee06d98abf4 · inbound

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning cites this paper.

G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:40.483549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:40.483549Z digest=sha256:8d35f33c905a4c671dd2ddf27aee227ee285405e77d155a07f7bbd0cc4554fa9

Observation d027aadf-72d6-4160-bbc7-99f50ff05e08 · inbound

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective cites this paper.

A Survey on Vision-Language-Action Models: An Action Tokenization Perspective JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 262

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:08:35.299571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T14:08:34.893876Z digest=sha256:7c2885220889de390639f4c43de49d6a68c1ea6f3dfe8701cfd3cc0ecbf09a59

Observation 6e5074be-22d6-4b5e-ab8f-c71c5202bea1 · inbound

InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow cites this paper.

InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:12.517727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:12.517727Z digest=sha256:b69e8506c02ee643e7d83a5422b26ac918f2bbdb6406377305043f8b894a7b22

Observation abb013ec-6241-4b19-a271-d9093a54621a · inbound

Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic Reasoning cites this paper.

Society of Mind Meets Real-Time Strategy: A Hierarchical Multi-Agent Framework for Strategic Reasoning JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:40.726005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:40.726005Z digest=sha256:797c78b0d4e06c0cb33f48d872282d6fed2e856dd8ba96cdfe0af252523dc9f9

Observation 2eff03f3-e254-4d06-88d9-e3454eef8b1c · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.004983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:dfa0361be027769ca8ce9a21425fe2972acd1478249d906a9edb49692c57856e

Observation 4ce6403a-02bc-4462-967f-09a4643747aa · inbound

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents cites this paper.

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:08.031769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:57:36.038091Z digest=sha256:f08a6bc434b6308355443c78e7f6a20c2a4e6f7ecf5ea3939008efe4a79a3246

Observation 80d67920-a61c-4882-b702-b4cd4cfbe0ff · inbound

World2Minecraft: Occupancy-Driven Simulated Scenes Construction cites this paper.

World2Minecraft: Occupancy-Driven Simulated Scenes Construction JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.350784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T08:01:12.571480Z digest=sha256:6717e4a0e6e5f4137f55160e15254fb905dd2459dac3d120d7046b902652d82d

Observation 608bc25b-f1c0-4202-b1d9-9aacf3f88b88 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:02:05.542043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:b8d0a1a171521b8c2ea4ffa0021cb5134ddd745cc44e1cc2f8d490356857650c

Observation 4ea5cb6b-4c02-4bc8-b4cb-0dc5a16a3bc7 · inbound

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? cites this paper.

Do Vision-Language-Models show human-like logical problem-solving capability in point and click puzzle games? JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:07:08.760783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T01:58:39.476408Z digest=sha256:bf4afd262b7061c25bfbf2907f36e287063a0181f03fbc930183366414881105

Observation d81108ae-4df3-4b7d-a0d3-65b445c56b69 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:02:34.523053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:dfdab057fe5c47e7a14aa28246fb1abe94dc5ceb966e29da48b50ddb9b64998e