Pith. sign in

Paper Citation Record · LEDGER

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models

As of 16 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2608.07314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07314 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T10:29:55.365544Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22fd29ac-b83c-409e-8d90-27297fe18a37 · outbound

This paper cites PaLM-E: An embodied multimodal language model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models PaLM-E: An embodied multimodal language model,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.195731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.231185Z digest=sha256:bddb8dd049ba0898e66f219163f979421d36115795af08c91f93b29985b4b30a

Observation 3fa02969-224b-4744-a1ae-5da4257e0daa · outbound

This paper cites RT-1: Robotics transformer for real-world control at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-1: Robotics transformer for real-world control at scale,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.186044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.235069Z digest=sha256:9088db610739c9fef261dc794face9d448ad8546acc1de5642e5d43ae52272f4

Observation 29a27b32-044a-4c8a-903a-48e1d3ac7059 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.175689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.238673Z digest=sha256:9c2e654cff3a8637c70becfc6e9ebeff5dc306f726f5f1734a0d91ff36bc67f6

Observation ea4da3e9-37b1-4529-9654-dc7e9ac20716 · outbound

This paper cites Open X-Embodiment: Robotic learning datasets and RT-X models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Open X-Embodiment: Robotic learning datasets and RT-X models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.165594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.242091Z digest=sha256:71ae58c6fb574f6cf69d09ddf9c4fddd792e703f1216dd3375d6321b4c97d5c0

Observation 56230c39-9f3d-4b74-b647-a0e692aa832e · outbound

This paper cites Octo: An open-source generalist robot policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Octo: An open-source generalist robot policy,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.155897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.245579Z digest=sha256:2a3e12768d73ce8a5fb6df0f5f46bcb6a2b2630bd1ccf27a413efb471ac73171

Observation 12848eb0-1192-4339-b31f-2d29fe98b861 · outbound

This paper cites OpenVLA: An open-source vision- language-action model,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models OpenVLA: An open-source vision- language-action model,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.147186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.248925Z digest=sha256:da726c4768e5f71758c05ce062eeeadc4aa4eaa43552d2e04c7aa574b925bba7

Observation bf801935-81ea-416f-8f26-aec8e029f257 · outbound

This paper cites Fine-tuning vision-language-action models: Optimizing speed and success,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Fine-tuning vision-language-action models: Optimizing speed and success,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.138028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.252088Z digest=sha256:234ce31a577bac555e92af39a3b4278d773c64f0287359a954a425af0a1c3dcd

Observation 0f01f29c-4545-49d7-95f6-b4fc8b577c71 · outbound

This paper cites A reduction of imitation learning and structured prediction to no-regret online learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models A reduction of imitation learning and structured prediction to no-regret online learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.127108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.254959Z digest=sha256:5e43f930434dfd646ea1b5193fdc104dd9314ea584bf961e558bbde586b81c4a

Observation 5476859f-65e1-4d2b-adcf-a043a72076f6 · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.257629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.257629Z digest=sha256:a053b1ae9c95f5a65ea884e995928726bc097cfe87f60eeea7d05604de479b14

Observation f73d1b42-aa20-4d12-8ef7-f7e48bd7899f · outbound

This paper cites SimpleVLA-RL: Scaling VLA training via reinforcement learning,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models SimpleVLA-RL: Scaling VLA training via reinforcement learning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.116999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.260771Z digest=sha256:f2db4ef026069cfdacd52b38eb085b550ce036642ac57ce0d66bf68f5d00c768

Observation 34301185-637a-412e-ab6e-6c5c09477126 · outbound

This paper cites RL Token: Bootstrapping Online RL with Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL Token: Bootstrapping Online RL with Vision-Language-Action Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.263923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.263923Z digest=sha256:72ff5fb5c3779460fda2c8c5c20981f53d5bdfb084fd93df3f22910acc6a9348

Observation 9f211459-f700-4493-a8cb-9098f4574656 · outbound

This paper cites Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Knowledge insulating vision-language-action models: Train fast, run fast, generalize better,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.106712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.266959Z digest=sha256:73a7d8752be37651bf1dacd3352fc9c1ad2d3ef76ce1e6d5a664092006a8fb7f

Observation edfee9b4-2bab-447b-8495-e59005959870 · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.270093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.270093Z digest=sha256:d71fc3557ac5d685b835ad6e8874513fd67df8ce6d646e01e59863dd60a13502

Observation a5f798a9-de3c-464c-b2f5-7123f70a7771 · outbound

This paper cites Addressing function approximation error in actor-critic methods,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Addressing function approximation error in actor-critic methods,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.096260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.273617Z digest=sha256:4ec54e4044ad544afabe95ff7f41f9cedec4d51954740717199500d25593f24d

Observation 7668e1bc-5b3f-4710-82f3-f085ce17c84f · outbound

This paper cites BridgeData V2: A dataset for robot learning at scale,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models BridgeData V2: A dataset for robot learning at scale,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.086289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.276679Z digest=sha256:8e45a1f3f4df28edfa510cd6dc6d7783e4d9345e5c6a2d1788fc19ae86bf99f5

Observation 29ca5594-d6bd-4964-94df-c05bdf3801a0 · outbound

This paper cites Vision-language foundation models as effective robot imitators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Vision-language foundation models as effective robot imitators,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.279752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.279752Z digest=sha256:c0e0f4783bd05f0f0ee066cdc3e5ccaf6fac7fc7d7478c6570631538c4769867

Observation 4b7bdbb6-6c16-4379-8ffe-24ff1d0f6af1 · outbound

This paper cites Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Visa-flow: Accelerating robot skill learning via large-scale video semantic action flow,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.915698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.283014Z digest=sha256:14c154efe26049574bcb45d7c692711633f174dc693c14fc2f66106ba594b48f

Observation 6dadfcca-bde9-4ada-9e09-db4045287d65 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π 0: A vision-language-action flow model for general robot control,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.070341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.286220Z digest=sha256:097e3b1429b4e19e3a0bdc7c077e51cc279d59db024cfc565700e8e9e0994393

Observation 26618d8a-a725-4ff9-845f-2dafac8509ee · outbound

This paper cites π0.5: A vision-language-action model with open-world generalization,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models π0.5: A vision-language-action model with open-world generalization,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.060558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.289517Z digest=sha256:cdf4ca370ca956fcea6f052ef1c4e1b1d1ca180027804b8f9980daf315c01c8f

Observation 87607237-d798-43ab-8d8a-065063327d4b · outbound

This paper cites GR00T N1: An Open Foundation Model for Generalist Humanoid Robots.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.296092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.296092Z digest=sha256:5b2ba7347d2aa358e0bc4a92ee6568dae591dad313252b75ffccc7195da8d08f

Observation 3066997d-de2d-47fe-ade4-6bcea7246a0e · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Interactive Post-Training for Vision-Language-Action Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.299516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.299516Z digest=sha256:de2fbe743946466b2ea2dff7c7f1f0ed0a825738891fd7de3952e20954345cd0

Observation eec1aa47-6d89-4fbe-bd06-e161450369d4 · outbound

This paper cites VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models VLA-RFT: Vision-language-action reinforcement fine-tuning with verified rewards in world simulators,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.040569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.302904Z digest=sha256:58d1c7fa4bee15448c6d3647acfed5b92b48dde85166cb9ba1ee370cb82624da

Observation 8599fa7e-bda9-4196-9164-a095a7ad8298 · outbound

This paper cites ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models ConRFT: A reinforced fine-tuning method for VLA models via consistency policy,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.030298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.309572Z digest=sha256:518242de314eabea172c94ac37474ca98429cd7d0c52a2d81e834480afde2816

Observation ef126e3e-aeae-4b5e-8752-2cedf4b0b831 · outbound

This paper cites $\pi^{*}_{0.6}$: a VLA That Learns From Experience.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models $\pi^{*}_{0.6}$: a VLA That Learns From Experience

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.312519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.312519Z digest=sha256:e5f8ae9cede005f05abd747cb6a759ad68cf3b05f4b94a477831ac601ce9a88a

Observation 6eb70623-a93a-45c2-8578-e2e368f4b264 · outbound

This paper cites Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Acting while understanding: Asynchronous semantic-action decoupling for real-time vision-language-action models,

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-10T10:29:55.660050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.316233Z digest=sha256:dc14e9ac4fde28face56ab9c0769d4f19a562a14d9e3dea2ea57a35bd05a2295

Observation bff91253-b426-414e-92fa-ed83fc42e6f1 · outbound

This paper cites DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models DAM-VLA: Decoupled Asynchronous Multimodal Vision Language Action model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.319520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.319520Z digest=sha256:638a0e077b1171d3f26c451c233c0380ea4f7f0af2d627a91b391b079850979b

Observation 1595bda8-f98a-4546-a9fd-5b91cd3c17be · outbound

This paper cites RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models RL-VLA$^3$: A Flexible and Asynchronous Reinforcement Learning Framework for VLA Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.322806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.322806Z digest=sha256:39c16d2e0df10818578445cdc02a99354606e40930010b1a1096682b521b831e

Observation c9190653-2b3a-4c89-be44-49d47b8d0d05 · outbound

This paper cites FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models FLOWER: Democratizing generalist robot policies with efficient vision-language-flow models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.326214Z digest=sha256:9e2d622563600cbde0e5a502e25942680a13b209b0f0873b699571f0dcb20024

Observation df257f01-d68e-4020-86cb-20df0eadc475 · outbound

This paper cites CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:56.010337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.329247Z digest=sha256:57b342079af0f9de9a011dbd15784cf1208b56e32ff30528a381ea3400b8ecfb

Observation 070e97bc-b8ca-414e-b4a4-7b62c0b62d05 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.332304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.332304Z digest=sha256:6f288ceefec85066f11f5985d5b3292e3d26d007ac9562ac6c978be0b974d3ab

Observation c0d5ca24-f249-4b16-aee9-1bbd471e85db · outbound

This paper cites Unified Vision-Language-Action Model.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unified Vision-Language-Action Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.335648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.335648Z digest=sha256:15cd52c96f223e881ba55f90ffd0ce21cc49573540bc8ad80e1b873fef47ebc5

Observation bcf3a1ed-5608-4cf1-85df-33d4ca959e10 · outbound

This paper cites Disentangled robot learning via separate forward and inverse dynamics pretraining,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Disentangled robot learning via separate forward and inverse dynamics pretraining,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.999192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.339006Z digest=sha256:249d11a251178926cc92416185d8536e25cdd3b96b65dafc671ba231df765249

Observation 494b723d-3fc4-4c09-be34-5b9b3e44242f · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.342197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.342197Z digest=sha256:9d62b1c36134a39199dac858afc585d396473c3a5e504cc89d29f8f30d46b400

Observation d3f1a906-5701-4748-890a-26e4d2cd3052 · outbound

This paper cites 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models 3D Diffuser Actor: Policy Diffusion with 3D Scene Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.345462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.345462Z digest=sha256:6c63a873bedf54a5de70b95c144c08ada84fbbe71f645b1e8de1d538a99df672

Observation 151c7e52-33d6-4715-9660-055547a9ecc6 · outbound

This paper cites Closed-loop visuomotor control with generative expectation for robotic manipulation,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Closed-loop visuomotor control with generative expectation for robotic manipulation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.988461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.348290Z digest=sha256:a8a2164f8f1c898271dbb51200f8a20f2edf9c168f7a38cfcddf48c4bee4705d

Observation 57f9af1d-d3b1-49dc-8040-fe607b8cfead · outbound

This paper cites Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Towards Synergistic, Generalized, and Efficient Dual-System for Robotic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.351046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.351046Z digest=sha256:95e2c2577017a18bdf50cc2f2cbf39d965e507a83f2adb34b5bf12f3e3cdd73d

Observation 329ece0c-ce06-44bb-bb75-991b50ca45d5 · outbound

This paper cites UP-VLA: A unified understanding and prediction model for embodied agent,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A unified understanding and prediction model for embodied agent,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.977851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.353821Z digest=sha256:0892a63f9eff07cc06d4c0b6b129df5b9042477e2e41def2a698f1c41130f13e

Observation 995e8f22-1a8e-46e0-84bf-941c89e41238 · outbound

This paper cites What matters in building vision-language- action models for generalist robots,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models What matters in building vision-language- action models for generalist robots,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.966734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.359366Z digest=sha256:7e4296dffcbb3fe84a4cf570dba725d21a5703a4185128930fa5f16f07c0ee2b

Observation b1c35929-4741-4869-a527-3e1902b19ff7 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.362007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.362007Z digest=sha256:374b2427d1e9ebc814db6c276334c946882d3ef848413356851cc42b1a76bb51

Observation 76e909f4-f93f-4e90-991a-6a937ec19a6b · outbound

This paper cites UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.356467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.356467Z digest=sha256:b98b7121b93ebe891a8bf93e9f49cd907132cc4ff5123a75f219f5dff0663268

Observation 07711bc4-1625-48fd-b1e5-c6201cd13b93 · outbound

This paper cites Video prediction policy: A generalist robot policy with predictive visual representations,.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Video prediction policy: A generalist robot policy with predictive visual representations,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T10:29:55.955746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.365544Z digest=sha256:28313cd1e75d7aad58f954da2911d76810dfa3545d5e4e4081df3c553b77b211

Observation 7294401d-902b-4fdd-bf7c-f2e75b9fb205 · outbound

This paper cites an unresolved cited work.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Unresolved cited work

Reference 305

Resolution
unresolved
raw_fallback, observed 2026-08-10T10:29:56.049969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T10:29:55.292911Z digest=sha256:b641e27dd6aa11b847a85b7d523fd4ca4a2a4fd149b3ace7647840cfbfc9744a

Observation c57bf128-f41f-4721-82f8-47351af4de48 · outbound

This paper cites Available: https://arxiv.org/abs/2510.00406.

TEMPO: Semantic-Action Decoupled RL Post-Training for Vision-Language-Action Models Available: https://arxiv.org/abs/2510.00406

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T10:29:55.306328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:29:55.306328Z digest=sha256:eabd714c645034fcc6089b414616781bb41fe3fd8f23f8ad9898fff5e4201f2b

Pith citing papers

No inbound Pith citation observations are available.