Pith. sign in

Paper Citation Record · LEDGER

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2509.04063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04063 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:30:23.710832Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:51:28.966906Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T12:43:17.233166Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e864b4a-3641-41f4-a30f-edb30d1a5381 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.444500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.444500Z digest=sha256:f212f5905f49a0e5157a08c78d599eb4bafdc2cb6f70ddf0f93312140e0620e8

Observation 22c3cd8a-a31b-4e15-af7b-55d50703bf4e · outbound

This paper cites write newline.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.451025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.451025Z digest=sha256:2a5cc3027a67189a1224a6c62e80c6f4b9dba8ee413a1eb040412abc7da7f71e

Observation 27ed1442-1a51-4744-899a-bc5d90be4341 · outbound

This paper cites B.; Jaakkola, T.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models B.; Jaakkola, T

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:30:24.520293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.457237Z digest=sha256:e139a01900e61dae229e5fea7c29a553adddad6577def4615cc2789499263594

Observation 57708748-ee58-4c4c-af3c-4baef8fcf1a6 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.463943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.463943Z digest=sha256:461ed69ccc77a9e332d3ba60373e6b8ac2fa787a90cb4c93440b3e314c6896ca

Observation 49033faf-97a7-4f53-9d63-22125454ec02 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.469543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.469543Z digest=sha256:5ce1def0e50d676a24971595fd59dbfda0125ac541f8c867f1659cfe86eb5a76

Observation aebd0cbe-21f0-4217-8ae9-d1bdccd21934 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.475841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.475841Z digest=sha256:14adef8eed2a146fa062e6c59e7ff45dda1813eef37a8c09b10a46ce5946f2a6

Observation fa1a897c-afaf-4e44-9a1c-d5a82059ee22 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.481852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.481852Z digest=sha256:d5c5db914bfe9563c7beda377548501450657c46ca633f1b9747cb46d11e7b55

Observation d63c2495-1ddf-4d7f-88f7-592d5332d574 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.488324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.488324Z digest=sha256:36e9482e0dcd4694852db19d377f603a0b5a785067c95a7f6d4769d0dd6786bd

Observation c214714c-c744-48ab-87ce-77d35e271900 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.490508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.495241Z digest=sha256:003b3a2c67e10c5328b721b23bf8c97585fbaf41a1f12bdb68cd579a49d2c9ff

Observation c308da86-c944-4c17-bc8f-eba9bc4d20ae · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.471335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.500979Z digest=sha256:3a1d56f5538fa0cf7f3a7cad9c6dc082ef5dcc38704edf70e112b467b9e82b1a

Observation bb9929da-5aa4-48c0-abff-d109b0443be8 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.454442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.508419Z digest=sha256:c15ada1d2c0a282c1ea30754ffc091b9dedf2431ea0c8b4c9227067c45dbfe0c

Observation dfe423ff-3d87-449c-a8c7-b60679f9a6a8 · outbound

This paper cites Neural Ordinary Differential Equations.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Neural Ordinary Differential Equations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.515915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.515915Z digest=sha256:8e6522540713ecefadf3bccc746217a16a79dd86ec07d89ee9cfd14bb7554c31

Observation f3e1c7e4-be1d-42dd-a51d-db9aa39d679a · outbound

This paper cites FDPP: Fine-tune Diffusion Policy with Human Preference.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models FDPP: Fine-tune Diffusion Policy with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.521640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.521640Z digest=sha256:79a0dfa2eabdc05bbdaa6ac244387ccf9af2357d4440ec7ab3985cc0871fafc4

Observation 7dace495-c2da-4461-814f-4faa88315a76 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.437323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.527844Z digest=sha256:754e2551c071faa16e1207c81a3534f6e6f56318b915d6cde1fdc78b4021348b

Observation f6f29c4d-4779-4f8d-9a3b-b6b65dca152f · outbound

This paper cites S.; Lynch, C.; Chowdhery, A.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; et al.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models S.; Lynch, C.; Chowdhery, A.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:30:24.419516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.533520Z digest=sha256:840058534626aee86ab5d0163411b1b7ffc7777e807e84d12dfe53b9c6ec1efa

Observation 35ba3196-ca77-4c1e-8e83-1549cb75e2ac · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.538701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.538701Z digest=sha256:344d679afb9aa320f7ee601b9c87468a858bae2bb9cfdcf59a01270ddb8a9c7e

Observation 06c6519d-9fdc-4bb8-9e5e-546e23baea4f · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Denoising Diffusion Probabilistic Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.546728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.546728Z digest=sha256:998258737b54543b9493453feddef3f922433d16d5d2b7277c7419f91a082969

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:1a1ed42b514ba1ad3a840e1d417c9e3e5051b3894cc4335b435224f179ab2256

Observation 4f31f06f-4b50-42b9-ad72-780caa8b1bbb · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.401852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.558652Z digest=sha256:ecd8911a10c158759c5f3d39fd7684acfd52491f70650c98061e1ddc82cc3ebf

Observation a8d857f0-8d2b-45f4-a6b8-b9b67151b6f9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.563766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.563766Z digest=sha256:fe73c1ae5f1e0d627298112e1619573598913015bd021d71f066d300b06ff808

Observation e284bcf8-56be-4f37-b6e2-3a5bd1de3fd5 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.382885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.570311Z digest=sha256:85449886684e8b2db38cca27a64d7fff40abe1a19e8bb3ece605d6404a4930ee

Observation 79743d51-d3a7-4fd5-a3a7-23ca43a66281 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.576111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.576111Z digest=sha256:c9bfb4ccc4bd96ca5abb39835980d9ba33235885899c9663641577374f768e2d

Observation fe88a4da-8409-4ba1-9ecd-04a26d19ce36 · outbound

This paper cites Flow Matching for Generative Modeling.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.582107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.582107Z digest=sha256:b3203fa920822401192ec6d6c0ebe7981b261f7c8e2089e159128203151a851f

Observation 1bdee1ee-8647-4c50-93e4-15dd49ab6174 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.587313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.587313Z digest=sha256:6d806a695fee630f4beb36682a154009a2861c869337f5a78e0138bf4aab19ea

Observation 2a368fcf-5caa-4743-9d43-14b625537fab · outbound

This paper cites Decoupled Weight Decay Regularization.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Decoupled Weight Decay Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.593021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.593021Z digest=sha256:41e20caf1a36e7a747bc33135096e10e7aec55e76a35d3860cdffabdc9578335

Observation 6e05dbf2-4666-4a20-870d-2f28538071a6 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.361360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.598619Z digest=sha256:d246c9f2208c5488702394fd62a0767025e4cd97eac0034da50b5d5d9c1f4833

Observation 7cc1bd6a-b221-4f54-b7e7-2c0543864e5a · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.605483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.605483Z digest=sha256:407a9d05328189d8cde8ebe9dc097211e2072740b4fd901324dc979c66b8b653

Observation cdd8c86c-0b3b-4f1a-af6f-f9b4914be986 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.611023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.611023Z digest=sha256:93e69cef51e7ee538b9decc6435ec0682458941f7b6d697df7970d3f4b268688

Observation 2f15b711-1537-4c34-acf5-e1a2688baef3 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.343397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.616193Z digest=sha256:c8660043b88500c967eedddd023761fa1ca1d6d5abad7caadd8f764eb585c971

Observation f81a2e7d-54e7-48a4-a8d8-e0ffe599cf5c · outbound

This paper cites Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.621321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.621321Z digest=sha256:103e7ff1d0c5b21ed8d81d7af702c93aa2a6e54d49a735e90db3d0faf1360eae

Observation 536ba1d1-401a-476c-bf6a-ebd0758db317 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.326271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.626576Z digest=sha256:ca62bb45d21feb19eb89e55414a4a04897b760af5fd72308176a5e424b8a6bba

Observation ffee3e7b-3f88-4282-9ee9-d2b050e62ed1 · outbound

This paper cites Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.635769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.635769Z digest=sha256:38dc5f5df6901f8d37d1a65b69f265de35f64dea9f0ede8e57be7376178d626a

Observation d44e2bf0-3e45-4016-bd0e-0320436923be · outbound

This paper cites Proximal Policy Optimization Algorithms.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.641166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.641166Z digest=sha256:b472a1ecd70a5a829a535afe18e267d6b02e95e243082232cde917c2505c6cdc

Observation 8bc3f0bd-77b6-4e4c-9cbe-07689334638b · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Interactive Post-Training for Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.646643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.646643Z digest=sha256:7d751f8da91d0ae47e00ab41e77e8f5bd9ed743ae8c84f24850e13219f9ce1ef

Observation 63d5175f-027b-465e-8d22-c3e5fbc4fa6b · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Octo: An Open-Source Generalist Robot Policy

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.652369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.652369Z digest=sha256:4d943654eaaecc594d44b49e93d1eda70ea67c715ee91ce9c3183b02017acaf6

Observation b8a9a791-599a-409c-89db-a0c000554697 · outbound

This paper cites J.; and Zhou, M.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models J.; and Zhou, M

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:30:24.308875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.658056Z digest=sha256:8961c5268824ca595299558ea75021914862d2333e1d599aa6f872d036cb4dfb

Observation dd6f62e4-7bcf-4126-aae7-ce8dc1fdf77f · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.289673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.664002Z digest=sha256:0f10daead86704040e2b1327f1dc523d0b896f1afae0703c8ebfdd843959da80

Observation c66a75f6-4f55-462d-b786-087c9b8bb200 · outbound

This paper cites ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.670740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.670740Z digest=sha256:d11f180c44f4ec99e07256d67a60425d2b6e04863ba50def018b521018050955

Observation 94d2adb0-a693-488f-9d59-14070dd5b872 · outbound

This paper cites Energy-Weighted Flow Matching for Offline Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Energy-Weighted Flow Matching for Offline Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.678531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.678531Z digest=sha256:9ee653f1b58179faab19a7f010ba2e7c32dcad0a38f1d1d219148e23e7f8b90c

Observation 02898c68-a010-4aa3-adf2-0cfe4243bf80 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.688049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.688049Z digest=sha256:32b2cc313a34f431850007c278e47d16b52bcb97f1cbd223ee19eaaa0e42f52c

Observation 8630cad3-afcf-41be-86f7-535f15cd85e6 · outbound

This paper cites MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.695864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.695864Z digest=sha256:f4890f79eec3fbaaf7a6d93c43e64bd75f1669948615b0bbaf322814ec474990

Observation 6ef7b8c5-a478-446f-a696-a7c7126604be · outbound

This paper cites Guided Flows for Generative Modeling and Decision Making.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Guided Flows for Generative Modeling and Decision Making

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.703140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.703140Z digest=sha256:dd5d57f662cb405f9ba6ea1333635c2580f0d1965d27d3237cfce810bd5b009f

Observation 5f98f877-341c-4bed-be13-e0f6251079c8 · outbound

This paper cites Reinformer: Max-Return Sequence Modeling for Offline RL.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Reinformer: Max-Return Sequence Modeling for Offline RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.710832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.710832Z digest=sha256:10495b1dc075e41b595bd1ec8010184d39f3099db6577392c94d99321b0a63fc

Pith citing papers

Observation 0515b405-0150-47b2-86a7-06e4b120a5ca · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:31:02.908493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:ef111069477740bf1d4101afd4a6c272d68e53c1ae2e733c7c89eafdd9a21110

Observation 523fe429-f1b8-4cb6-8c19-09b961db3bf4 · inbound

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing cites this paper.

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:28.966906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:28.966906Z digest=sha256:727fd1d39a2681ca46a062293f2c9b9c7d7117cd71bf55971ca19b070d7366f1

Observation 6ef27ad8-2026-4fda-934b-461fa5a9a422 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.235496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:b11e493cc614b7a576dec748bd77029c278c7875803810962b97ed5a527ea22a