Pith. sign in

Paper Citation Record · LEDGER

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

As of 16 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07314 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:29:55.365544Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22fd29ac-b83c-409e-8d90-27297fe18a37 · outbound

This paper cites PaLM-E: An embodied multimodal language model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models PaLM-E: An embodied multimodal language model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.195731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.231185Z digest=sha256:414088f0ed8b6999bb06b8d823e71ce4559f1518d7ff0ba981fed62dcdd51a85

Observation 3fa02969-224b-4744-a1ae-5da4257e0daa · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.186044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.235069Z digest=sha256:f4e11d6ce9c49521e3ae28efbfdda07466be893ce055bd8252df0afa65aeea3b

Observation 29a27b32-044a-4c8a-903a-48e1d3ac7059 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.175689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.238673Z digest=sha256:f7cd83211836276857fe0ff144f7d7c2f9930cdbc1e5c9a416546026756dc668

Observation ea4da3e9-37b1-4529-9654-dc7e9ac20716 · outbound

This paper cites Open X-Embodiment: Robotic learning datasets and RT-X models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Open X-Embodiment: Robotic learning datasets and RT-X models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.165594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.242091Z digest=sha256:64139d7ba67d6672535e7f9e9008aa7d67a9cfd16a31b4ba06ac8da3b67cb3dc

Observation 56230c39-9f3d-4b74-b647-a0e692aa832e · outbound

This paper cites Octo: An open-source generalist robot policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Octo: An open-source generalist robot policy,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.155897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.245579Z digest=sha256:4eb2b188202f888b6922f00602d3be8454c62535a5041f7a2273ce51a2ed2aaa

Observation 12848eb0-1192-4339-b31f-2d29fe98b861 · outbound

This paper cites OpenVLA: An open-source vision- language-action model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models OpenVLA: An open-source vision- language-action model,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.147186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.248925Z digest=sha256:f0c5f59140e024af960f6e2e41ba43ffad8869604eb79ba9627f645734f3e4a6

Observation bf801935-81ea-416f-8f26-aec8e029f257 · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Fine-tuning vision-language-action models: Optimizing speed and success,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.138028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.252088Z digest=sha256:a4dcfa8ff2bdfa5086538c799b60e722a5bb570eb6ec4822a2e3c7eec510c7db

Observation 0f01f29c-4545-49d7-95f6-b4fc8b577c71 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.127108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.254959Z digest=sha256:7287c44002a1d7d5ca3e781900a747a93fd68a0aff94c3de3c81b4e0da6c40f6

Observation 5476859f-65e1-4d2b-adcf-a043a72076f6 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.257629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.257629Z digest=sha256:c4f84e71f49e9139936accde37b052b3bb66c92f5e81fd669e9882de84259bbb

Observation f73d1b42-aa20-4d12-8ef7-f7e48bd7899f · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.116999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.260771Z digest=sha256:7af8e39f776507a1bc716177d3dc9a2944acc62133e0565bd899d5ee96772e81

Observation 34301185-637a-412e-ab6e-6c5c09477126 · outbound

This paper cites RL Token: Bootstrapping Online RL with Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.263923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.263923Z digest=sha256:f6be4e1cc2813d9a234585d3ddbd7136a24335a4bb9316dcb7c23d0b0e28a637

Observation 9f211459-f700-4493-a8cb-9098f4574656 · outbound

This paper cites Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.106712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.266959Z digest=sha256:4fd6acddd4ae92039c7b0720ed76a92c5de0634ad53505d1f97dcec5af9a0d7b

Observation edfee9b4-2bab-447b-8495-e59005959870 · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.270093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.270093Z digest=sha256:b9f8938ba55e845984abff6c311f80c90a27a2ee3b253af77c843ed91696df29

Observation a5f798a9-de3c-464c-b2f5-7123f70a7771 · outbound

This paper cites Addressing function approximation error in actor-critic methods,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Addressing function approximation error in actor-critic methods,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.096260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.273617Z digest=sha256:e5d79315dc53a533eebbb7a042c6792cb1db0cb7f4da118cb66e35ab778fe28b

Observation 7668e1bc-5b3f-4710-82f3-f085ce17c84f · outbound

This paper cites BridgeData V2: A dataset for robot learning at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models BridgeData V2: A dataset for robot learning at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.086289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.276679Z digest=sha256:053aa2ff078ed8f5d7ac1cef94646c7eb45581c0c976a8a7dcafb731c96ce546

Observation 29ca5594-d6bd-4964-94df-c05bdf3801a0 · outbound

This paper cites Vision-language foundation models as effective robot imitators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Vision-language foundation models as effective robot imitators,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.279752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.279752Z digest=sha256:9724d1e97fe66b29a8d16dbdf7e358123f30463052c4bf30c8e635755993aafb

Observation 4b7bdbb6-6c16-4379-8ffe-24ff1d0f6af1 · outbound

This paper cites Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.915698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.283014Z digest=sha256:5d2982d3cb916b8317b306907443b380d9cd115b80aab9676f5dc58817036e88

Observation 6dadfcca-bde9-4ada-9e09-db4045287d65 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π 0: A vision-language-action flow model for general robot control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.070341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.286220Z digest=sha256:d30f86cf8e7d2cc35a25aa6c30e4e511cf7b6925566cb3105d68e3b18f50f722

Observation 26618d8a-a725-4ff9-845f-2dafac8509ee · outbound

This paper cites π0.5: A vision-language-action model with open-world generalization,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π0.5: A vision-language-action model with open-world generalization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.060558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.289517Z digest=sha256:7f49a3cf0324155545773cc9fa88cdf8a1d3a9ef2c03d0bbff21f294823e713b

Observation 87607237-d798-43ab-8d8a-065063327d4b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.296092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.296092Z digest=sha256:336a2a4f426c86169391d1075b15e1e98fd60435c08e050016df2851d3dd8ae1

Observation 3066997d-de2d-47fe-ade4-6bcea7246a0e · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.299516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.299516Z digest=sha256:62584f9857c8e1ea292b4dbcf120d496d41b498e7628b8e9e96a998101e63792

Observation eec1aa47-6d89-4fbe-bd06-e161450369d4 · outbound

This paper cites VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.040569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.302904Z digest=sha256:050e2cb34a8bcf9a2fe9c57b88f444fd2b6990981e1cef0eaeaaa517c7ab5580

Observation 8599fa7e-bda9-4196-9164-a095a7ad8298 · outbound

This paper cites ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.030298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.309572Z digest=sha256:3a5d1537c100e88756128e36da21564c13badc7344833afffe9fdfeae7c5e9a8

Observation ef126e3e-aeae-4b5e-8752-2cedf4b0b831 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.312519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.312519Z digest=sha256:849fb8966445935639acfdf8c4bcb8f2db91ce6cf20557720a99d0260a443ca3

Observation 6eb70623-a93a-45c2-8578-e2e368f4b264 · outbound

This paper cites Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.660050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.316233Z digest=sha256:0d26759d27120d4943a1457dea24c18147b5cd8d17f1bd1bc3d110ba94ee52b4

Observation bff91253-b426-414e-92fa-ed83fc42e6f1 · outbound

This paper cites DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.319520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.319520Z digest=sha256:e71803d38b87a89fd11190fb14fecc771df60c7eef1834e307924e92a30f275d

Observation 1595bda8-f98a-4546-a9fd-5b91cd3c17be · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.322806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.322806Z digest=sha256:960a962aea51833b671c701df6b81b931dbd3a630ab2889bdd7604d207e8acdd

Observation c9190653-2b3a-4c89-be44-49d47b8d0d05 · outbound

This paper cites FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.326214Z digest=sha256:497f4fe0047a306c66838221ae04f9cf90bc370c3852cccbb0a29eb363061c6c

Observation df257f01-d68e-4020-86cb-20df0eadc475 · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.010337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.329247Z digest=sha256:9b53aba1129be59df91a59d08b63f7e824deae74eaa17efdb8c1e1d3d4ccf04d

Observation 070e97bc-b8ca-414e-b4a4-7b62c0b62d05 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.332304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.332304Z digest=sha256:3e7ec311dd38abbf66e6b2d48355ff1ff83e59a5645fadf1c0d1b67602f39777

Observation c0d5ca24-f249-4b16-aee9-1bbd471e85db · outbound

This paper cites Unified Vision-Language-Action Model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.335648Z digest=sha256:ec046fb3f00da328125ed66e932e2e18e9cfab00af89ca4e8ce8ae12be7c058b

Observation bcf3a1ed-5608-4cf1-85df-33d4ca959e10 · outbound

This paper cites Disentangled robot learning via separate forward and inverse dynamics pretraining,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Disentangled robot learning via separate forward and inverse dynamics pretraining,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.999192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.339006Z digest=sha256:04257887660cd834e4730b5fe76888cb1094f651c3f0a7bdd27dc3ac20c0b81b

Observation 494b723d-3fc4-4c09-be34-5b9b3e44242f · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.342197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.342197Z digest=sha256:03669f43e81459e87fff91b852c7b37b8ea135bdb4b497f741c5bc811660aa61

Observation d3f1a906-5701-4748-890a-26e4d2cd3052 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.345462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.345462Z digest=sha256:f73c618aead4f12010bf9ab966d4c2c424b04c03f57173119fe1b6c9c0750c15

Observation 151c7e52-33d6-4715-9660-055547a9ecc6 · outbound

This paper cites Closed-loop visuomotor control with generative expectation for robotic manipulation,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Closed-loop visuomotor control with generative expectation for robotic manipulation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.988461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.348290Z digest=sha256:668945eb9cd293bdf5c08190ab4611a433df5d72d6bc58c60a3adfb7f0830bdd

Observation 57f9af1d-d3b1-49dc-8040-fe607b8cfead · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.351046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.351046Z digest=sha256:7274740e0dec0a98bdb00c4181fd9755ca608deb5ecdfa09fc87106c3d309bac

Observation 329ece0c-ce06-44bb-bb75-991b50ca45d5 · outbound

This paper cites UP-VLA: A unified understanding and prediction model for embodied agent,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A unified understanding and prediction model for embodied agent,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.977851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.353821Z digest=sha256:8f9459da685fcb028c5ff0e2c9c62d7a9466e84cdec0ef015d51373a452f5078

Observation 995e8f22-1a8e-46e0-84bf-941c89e41238 · outbound

This paper cites What matters in building vision-language- action models for generalist robots,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models What matters in building vision-language- action models for generalist robots,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.966734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.359366Z digest=sha256:b5acb2663e241f92d5a03371c6c16d453c2bcc3eba7fc9ab58ade4f1188ab964

Observation b1c35929-4741-4869-a527-3e1902b19ff7 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.362007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.362007Z digest=sha256:1dc51f922dbd586266ccb447bfccb1de008ae22db3ac47f93e5b0e17089f0739

Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.356467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.356467Z digest=sha256:9b5b6a0ba82a8c7e2b9540576974e10d37955507db2c474d752747764ab5f7a9

Observation 07711bc4-1625-48fd-b1e5-c6201cd13b93 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Video prediction policy: A generalist robot policy with predictive visual representations,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.955746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.365544Z digest=sha256:792bf4e75e4186d85d33cec0996a43484c02962a60bb56306c011a4f336f6e2b

Observation 7294401d-902b-4fdd-bf7c-f2e75b9fb205 · outbound

This paper cites an unresolved cited work.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unresolved cited work

Reference 305

Resolution
unresolved
raw_fallback, observed 2026-08-10T10:29:56.049969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T10:29:55.292911Z digest=sha256:811cf8601cd3fc93482ff7042ff7a5b7194b1a2e255225ad08b211153ffc8553

Observation c57bf128-f41f-4721-82f8-47351af4de48 · outbound

This paper cites Available: https://arxiv.org/abs/2510.00406.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Available: https://arxiv.org/abs/2510.00406

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.306328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.306328Z digest=sha256:a4e09f536bf908abd19b9244dfd37ca6a884b9233a7bc60330d94f9724307537

Pith citing papers

No inbound Pith citation observations are available.