Pith. sign in

Paper Citation Record · LEDGER

From Foundation to Application: Improving VLA Models in Practice

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 7 inbound Pith citation observations for arXiv:2607.06403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06403 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T07:06:13.473493Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:16:39.641954Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T14:52:36.159020Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact31
  • verified fuzzy14
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5cfc9951-4d25-4d42-a0ba-7eb2e0803ad2 · outbound

This paper cites Cosmos 3: Omnimodal World Models for Physical AI.

From Foundation to Application: Improving VLA Models in Practice Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.485926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:9b8a22615e8fa082b952a81e00f66ab0cc43bad511d067b6d4551fb34a26c4c9

Observation ad679c6b-c74c-46c5-8811-c36109dc0cf5 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

From Foundation to Application: Improving VLA Models in Practice V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.493399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:804591f3a7d0289e3a1d5d5bebe735b4574aa4e88bf6d38c04aaf0dd9c109953

Observation 003e4f82-c2c8-4c38-9672-109e2ab194a2 · outbound

This paper cites HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation.

From Foundation to Application: Improving VLA Models in Practice HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.498061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:e2c93ba6a71fa24ddfc18f4c3175a0a244c30d4019d7c66e97a69b9103b32e8e

Observation f74c6e13-d353-4e0f-8184-eb646ba6a0bc · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

From Foundation to Application: Improving VLA Models in Practice GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.502118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:b9f75ce30e4dc4088d93f45ec8ad2205243664a48187e284e0b518b93a5f31b6

Observation 86cab0f7-a598-4d4c-b632-021ef85ad981 · outbound

This paper cites an unresolved cited work.

From Foundation to Application: Improving VLA Models in Practice Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-08T07:24:43.712960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:ec90ca816577a77efa1fb4b4c3a67958c4b7ae65655d58201b99ad81b741ea7f

Observation dda130c9-d3f6-40bd-a14a-ac9bfca4c787 · outbound

This paper cites InProceedings of Robotics: Science and Systems.

From Foundation to Application: Improving VLA Models in Practice InProceedings of Robotics: Science and Systems

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.704491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c9c8d45d864b9e7cb9ee72dc22fb3bac1893fd1bf11f8ea210b2be7afeeab698

Observation 91d09b23-e93c-4f91-a230-f1db8a83125a · outbound

This paper cites arXiv preprint arXiv:2602.12684 (2026).

From Foundation to Application: Improving VLA Models in Practice arXiv preprint arXiv:2602.12684 (2026)

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.505934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f9eed1be5ef6b4b2bbe6e6d8317026b7ea0d16dd2e86ccfd9c09dcfbdbe20218

Observation 7d8a7ae8-e647-4482-9483-5308709382ad · outbound

This paper cites GR-3 Technical Report.

From Foundation to Application: Improving VLA Models in Practice GR-3 Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.482564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a23f5ed96732c6c560940ffe961277c6a651792d388b4c3e42e9228108ada986

Observation 84bce366-b3c6-4b47-aa2e-27f5c335d9db · outbound

This paper cites Lawam: Latent world action models for efficient dynamics-aware robot policies.arXiv preprint arXiv:2606.15768, 2026.

From Foundation to Application: Improving VLA Models in Practice Lawam: Latent world action models for efficient dynamics-aware robot policies.arXiv preprint arXiv:2606.15768, 2026

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.468230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c3df3f63f710a0eea0daa55194a9ae2927a8328f0ae263b624d42e8127fa4f46

Observation 49ec92c7-707e-4b92-9041-be7498122d83 · outbound

This paper cites ABot-M0.5: Unified Mobility-and-Manipulation World Action Model.

From Foundation to Application: Improving VLA Models in Practice ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.472368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:e994c8f44f407dae437d81ba9359761aa35528c080a86fe03071f4e1d9cd7947

Observation bc4f3dae-fafa-47c1-ba37-c76f365535c6 · outbound

This paper cites Dexworldmodel: Causal latent world modeling towards automated learning of embodied tasks, 2026.

From Foundation to Application: Improving VLA Models in Practice Dexworldmodel: Causal latent world modeling towards automated learning of embodied tasks, 2026

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.675975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:3b78b4bc121a66f7f9416bbf36a655e22f9030c0dd8e9559a2a414c244afa1d0

Observation 22113d00-53bb-49f0-bb94-23dbbf411775 · outbound

This paper cites HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies.

From Foundation to Application: Improving VLA Models in Practice HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-09T02:19:49.242114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c4183480fe4157789ba6d7647bd5f6912fb349beee9b1f6995309a212e39d9f0

Observation cc3a46fa-a757-4bb3-a01b-df17909688d3 · outbound

This paper cites Galaxea g0.5 technical report, 2026.

From Foundation to Application: Improving VLA Models in Practice Galaxea g0.5 technical report, 2026

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.710268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:191db30b47bff0d28b3d7b471b5546ed4bcde4c4108b8030c8d766ebbda9804d

Observation 853afccb-f2b6-4f39-8a32-ce3bea1f01f7 · outbound

This paper cites Galaxea Open-World Dataset and G0 Dual-System VLA Model.

From Foundation to Application: Improving VLA Models in Practice Galaxea Open-World Dataset and G0 Dual-System VLA Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.462253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:8f095e1e5e5892e69adbe68dbebdac1afb130dce9ded3c845e812872f3689327

Observation b0d968cf-f23c-4deb-b81d-c421f2bd1c1c · outbound

This paper cites RLDX-1 Technical Report.

From Foundation to Application: Improving VLA Models in Practice RLDX-1 Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.478990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:dc7ece82eeb7ac334ed9d392f94f4b727c97bff146d7baa4f030cbaf2aeff104

Observation 10e80ea4-bf94-4929-b266-8784b39cb6cf · outbound

This paper cites OpenVLA: An open-source vision-language-action model.

From Foundation to Application: Improving VLA Models in Practice OpenVLA: An open-source vision-language-action model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.681788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:73db0c7082889dbf45fc601212658ea29496cd43c5850b60eadefceffbb8570b

Observation fc61d89c-b1a4-4adc-ba11-40f28e2e2673 · outbound

This paper cites Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich manipulation.

From Foundation to Application: Improving VLA Models in Practice Forcevla2: Unleashing hybrid force-position control with force awareness for contact-rich manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.454088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:acd46892990dc37b9f9ee7f55ec16f11e64dba50814b57f446815cfc0d5c4ae0

Observation 414594b6-6e00-4fa7-923a-6c634c02b9fe · outbound

This paper cites HoloBrain-0 technical report.

From Foundation to Application: Improving VLA Models in Practice HoloBrain-0 technical report

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.438202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:311b255eb66d38509578a33b1b0c807334391dc5de35cb4037a27dc361196844

Observation 7a9673b0-6999-4e6f-9221-507e0d6c44ba · outbound

This paper cites DeepSeek-V3 Technical Report.

From Foundation to Application: Improving VLA Models in Practice DeepSeek-V3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.441802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:483fe831c59cf3142f57696fc8165ea6dd138b1dde6003ed5dac6170d0c1be6d

Observation c898bc55-29e6-4ff2-b04c-63542be34257 · outbound

This paper cites Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos.

From Foundation to Application: Improving VLA Models in Practice Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.445289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2cdfce18c6e45396a3ec14ac73b09e194fa1db2a81819d586adcd2d2c0e1e948

Observation 6c238041-d146-4740-8dc5-4dd34b9f904b · outbound

This paper cites Being-h0.

From Foundation to Application: Improving VLA Models in Practice Being-h0

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T07:14:45.449397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c5a8584fb6bf48b1230d82880534453218d25c541b98a0105586e764015ad73c

Observation 91bcbe7b-0b66-45da-a1f9-82fd7c9fa25c · outbound

This paper cites Being-H0.7: A Latent World-Action Model from Egocentric Videos.

From Foundation to Application: Improving VLA Models in Practice Being-H0.7: A Latent World-Action Model from Egocentric Videos

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.425343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:cfac0a10e07508af47b7f469dc60b9e8709e87da0def6c312b74ab1a1ced5d7a

Observation 9c7c5ceb-597e-468c-b36b-ef4f15c62d99 · outbound

This paper cites LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion.

From Foundation to Application: Improving VLA Models in Practice LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T07:14:45.429695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:4eafb68398d1dfec14827cfa506a2bd9e70cfb94442c1c200012a7bafecf9304

Observation 09d20cf5-02cc-40c1-b310-e721e2ea7bbf · outbound

This paper cites LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment.

From Foundation to Application: Improving VLA Models in Practice LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.433532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:e3bb537927f3d7316ed5b939a6f4dcb6c57ed80707816d71859600ee075be45c

Observation a7e6a948-fdb6-4cbc-9887-7506ede97bf2 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026.

From Foundation to Application: Improving VLA Models in Practice Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.693592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:b20e0bc46d729f9d3d9e061597f24263c6823946fd1dca88985643f975d0a7a7

Observation 200de719-b032-4001-99d9-62d85cf3582a · outbound

This paper cites an unresolved cited work.

From Foundation to Application: Improving VLA Models in Practice Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-08T07:24:43.695786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:620c0671021e217df090bcb027063e9dff3c8334e4db8f3a028400fc55bf041f

Observation a50f5f59-ac4b-4439-9906-daba83719e83 · outbound

This paper cites World guidance: World modeling in condition space for action generation.

From Foundation to Application: Improving VLA Models in Practice World guidance: World modeling in condition space for action generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.687497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:52156a5e4266a29fb760c058a487ab1ef2acce28579713b3324bca0bd53bcc18

Observation 1d7926b9-6a1e-44e7-b461-e6af4cac3c85 · outbound

This paper cites Masked depth modeling for spatial perception.

From Foundation to Application: Improving VLA Models in Practice Masked depth modeling for spatial perception

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.678764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:7e4f9b43aea6799e6db4210c25b38606f1f00be7bd1e03c43c2055c5b924be3b

Observation d86c13d0-8e17-4dfe-ac6a-dab771786a42 · outbound

This paper cites Towards human-like manipulation through RL-augmented teleoperation and mixture-of-dexterous-experts VLA.arXiv preprint arXiv:2603.08122, 2026.

From Foundation to Application: Improving VLA Models in Practice Towards human-like manipulation through RL-augmented teleoperation and mixture-of-dexterous-experts VLA.arXiv preprint arXiv:2603.08122, 2026

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.408586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:422a07d65cef72d756eeac061e8a6cc639d98a409e90e88f5522ef7cecec2db5

Observation f87599c4-14d9-4a46-97d3-26cbd1f2227d · outbound

This paper cites Gemini Robotics: Bringing AI into the Physical World.

From Foundation to Application: Improving VLA Models in Practice Gemini Robotics: Bringing AI into the Physical World

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.412977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:4363abfc36e112c2835b7cbed626e18ff2f0a89eb3c8d39b1fbd032fdf741b72

Observation e1815cef-e8cf-43ef-8c62-97586520cf08 · outbound

This paper cites Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026.

From Foundation to Application: Improving VLA Models in Practice Gen-1: Scaling embodied foundation models to mastery.Generalist AI Blog, 2026

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.701906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:71ffabc3764b1de9c3b1244118760abb529fa96c5c9764d8b7c592ba89f9db49

Observation 6fc84b0e-6424-425a-b3e2-68d61766908b · outbound

This paper cites GR00T N1.6: An improved open foundation model for generalist humanoid robots.

From Foundation to Application: Improving VLA Models in Practice GR00T N1.6: An improved open foundation model for generalist humanoid robots

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.684694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:853e82dd41354a5c311b16dc1b72a528a5eb8825d446366948f76ed1525ac5c7

Observation 9a260665-5e6e-4215-bd3a-d4184da1253e · outbound

This paper cites Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models.

From Foundation to Application: Improving VLA Models in Practice Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.707305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:f03bdf21a946fdcad1125d3182f6e4a26e747843d65660a352b86cf48012ba49

Observation e3815cc7-0ec4-4778-a075-b92499ecd99c · outbound

This paper cites Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments.

From Foundation to Application: Improving VLA Models in Practice Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.417284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:c56af1a0ecbde676f96364351a7b1225e6106da982bb2882b1390a4b520fb1b7

Observation dfc55af1-d17c-4f73-bb2c-ae5e12688dc7 · outbound

This paper cites Vision-centric activation and coordination for multimodal large language models.

From Foundation to Application: Improving VLA Models in Practice Vision-centric activation and coordination for multimodal large language models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.398070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a90e1879bf9b23a5606bff51e62205577351b4bd01f871b656e9b391e7f4cdd5

Observation e14e13a8-c8c0-411d-8d07-72518703617a · outbound

This paper cites The Great March 100: 100 detail-oriented tasks for evaluating embodied ai agents.

From Foundation to Application: Improving VLA Models in Practice The Great March 100: 100 detail-oriented tasks for evaluating embodied ai agents

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.698622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:b478dcc275ec1575c5d09d6e3772d8b87e2dff49babbf5e13ff1ea97161a79d2

Observation a33f30d5-2454-4287-bcc9-6475f0bd0426 · outbound

This paper cites Videorope: What makes for good video rotary position embedding? InInt.

From Foundation to Application: Improving VLA Models in Practice Videorope: What makes for good video rotary position embedding? InInt

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.690508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:732ad68b21613f41255c53b7d9c4c7ce52e39e90eb02680bef2252c5901a3fc4

Observation fad77dea-ea50-4f54-a5e8-29f3e5b9bfa3 · outbound

This paper cites A Pragmatic VLA Foundation Model.

From Foundation to Application: Improving VLA Models in Practice A Pragmatic VLA Foundation Model

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.387940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:02ed5bec1a48f7ef3e10a71823b655b1a2a7a5824fdc75d034f2852695d82ebe

Observation 826ef41b-3fe4-4b80-82aa-a6989304975a · outbound

This paper cites Hy-embodied-0.5-x: An enhanced embodied foundation model for real-world agents.

From Foundation to Application: Improving VLA Models in Practice Hy-embodied-0.5-x: An enhanced embodied foundation model for real-world agents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.718522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:db5f1c29b02b54e2c6e8cb60456624f249399d750917ce4bebc52205cd1b3375

Observation b24c01e9-c49b-49d3-9a91-7304df6f9778 · outbound

This paper cites OmniStream: Mastering perception, reconstruction and action in continuous streams.

From Foundation to Application: Improving VLA Models in Practice OmniStream: Mastering perception, reconstruction and action in continuous streams

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.393708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:ce9b2bc45923a058f99224053d19e36abfc0ef667240ad4dbc1399b9f69545e3

Observation 0151f29a-b4ac-4cd7-841c-c12b40ffbc18 · outbound

This paper cites Magma: A foundation model for multimodal AI agents.

From Foundation to Application: Improving VLA Models in Practice Magma: A foundation model for multimodal AI agents

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T07:24:43.715649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:391ee0525c784c779c3f1c82ec276155a87a738d561454bd397ffe6e68416767

Observation 233158c1-bc37-43e6-9c03-04ddb6cefe9f · outbound

This paper cites ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning.

From Foundation to Application: Improving VLA Models in Practice ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.381619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:a6a40b94c9180bed64dc6499d9e91fe4c7e07177542ba8a7d8e960c8c0bbdede

Observation 77d13554-2337-4648-a28b-1f0e9f6599ba · outbound

This paper cites StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems.

From Foundation to Application: Improving VLA Models in Practice StarVLA-$\alpha$: Reducing Complexity in Vision-Language-Action Systems

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.385142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:eb2f8258ad3ef970eca3dab743f335f598c0fca22a2fba0b91005fb3deb533a4

Observation 738073c5-f75c-40ba-aed7-3ee744efdefc · outbound

This paper cites Samoe- vla: A scene adaptive mixture-of-experts vision-language-action model for autonomous driving.

From Foundation to Application: Improving VLA Models in Practice Samoe- vla: A scene adaptive mixture-of-experts vision-language-action model for autonomous driving

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.403016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:d414f91dd3ffbe7b973190aa30f8c7f8d94752df5ccd1c62e84ad859a46d3e9f

Observation 7019218b-1805-44b6-9062-ccefc4d59f97 · outbound

This paper cites Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation.

From Foundation to Application: Improving VLA Models in Practice Forcevla: Enhancing vla models with a force-aware moe for contact-rich manipulation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.421603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:16cb045ec679311e3be12b0811127c7579f4c1545459327768a877b5df5d1c47

Observation 9d0f20ee-5153-47d9-8194-1471d3e15e7f · outbound

This paper cites Wall-OSS-0.5 Technical Report.

From Foundation to Application: Improving VLA Models in Practice Wall-OSS-0.5 Technical Report

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.374157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:d38fbb09439f0644c2d86c52ce41b6d7ff8c896462cd0e33316518f40d5d93a8

Observation b61eabe8-f1df-46c2-bb7a-665fb8187bde · outbound

This paper cites Igniting vlms toward the embodied space.

From Foundation to Application: Improving VLA Models in Practice Igniting vlms toward the embodied space

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.377496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:199eed281aa46b93bcf92e9ea7a34a3ab813c904761943366c5116f93311e36c

Observation b610e1b9-44f4-4f9c-867c-890dbb7e5f3b · outbound

This paper cites AtomicVLA: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026.

From Foundation to Application: Improving VLA Models in Practice AtomicVLA: Unlocking the potential of atomic skill learning in robots.arXiv preprint arXiv:2603.07648, 2026

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-08T07:14:45.370620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:2b657c14afe669701d1bcdfd2db3bd4573e3468b85be2a33d3c9eed277fa0651

Observation 5f4a89b6-e6f7-4190-8252-542136a69e98 · outbound

This paper cites GEM: Generative Supervision Helps Embodied Intelligence.

From Foundation to Application: Improving VLA Models in Practice GEM: Generative Supervision Helps Embodied Intelligence

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.366955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:0a7eb0eb2583b92688a38d8dad9985adea6ec44cbcc690a87ad13c3b47438ebf

Pith citing papers

Observation 98631932-0bc3-4b2c-b9cf-0519035d1d01 · inbound

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation cites this paper.

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation From Foundation to Application: Improving VLA Models in Practice

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T16:20:01.326319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:20:01.326319Z digest=sha256:0fe1d4eb36ef5c120d43fa334c06fdd01a2f3c1ced45b485785fd79f009d232e

Observation b9201e3b-071c-424e-9891-3f7e9cbdeee6 · inbound

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation cites this paper.

$N_0$-TWAM: Scaling Tactile-Native World-Action Model for Contact-Rich Manipulation From Foundation to Application: Improving VLA Models in Practice

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-30T12:42:18.596168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:42:18.596168Z digest=sha256:5a023952e1c490f6e8f3d4990efb4bc6c090347242cdaca6255aea28df255864

Observation a10fb45f-1def-4378-b5c1-d489d00bbf8c · inbound

Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA cites this paper.

Route by Kinematics, Act by Observation: Kinematics-Supervised Expert Routing in MoE-Augmented VLA From Foundation to Application: Improving VLA Models in Practice

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-30T20:30:33.298679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T20:30:33.298679Z digest=sha256:77ea4108663a36cc3c1844901d4c456399e68e0183b85e0fbf6ce663db426784

Observation dc7be8e0-8ee9-462e-97ca-e942cd233604 · inbound

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation cites this paper.

Learning Panorama-Aware VLA for Mobile Manipulation with Whole-Body Teleoperation From Foundation to Application: Improving VLA Models in Practice

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:54.173311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:39:54.173311Z digest=sha256:bfd8604d46420619f0414d49c5d7e10a267b08194b99c4695fc2a638f05d19ae

Observation d3e28936-f8d4-4800-9924-c778c4249614 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud From Foundation to Application: Improving VLA Models in Practice

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T14:52:36.161955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T14:52:35.400227Z digest=sha256:b6afde89628e4106ff7efabe3c1eca240c4e353893b03bba1f73e8aaa9de5ea7

Observation 0aaf8514-feb4-413c-919f-035b15ed4ac0 · inbound

FACT: Failure-Aware Causal Training for World-Action Models cites this paper.

FACT: Failure-Aware Causal Training for World-Action Models From Foundation to Application: Improving VLA Models in Practice

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:16:39.641954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:16:39.641954Z digest=sha256:7eb84185899d0dc0224221cc0060fb65009f000431f7edc8c8cf08e6aa4c21ae

Observation 50d7803a-337b-4a30-844b-00f571cee7df · inbound

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility cites this paper.

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility From Foundation to Application: Improving VLA Models in Practice

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:40:38.437224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:40:38.437224Z digest=sha256:c8d8e41f2bad8afea88d60564d9b69fd6377bd2b28f0aa892f3d44f0be002018