Pith. sign in

Paper Citation Record · LEDGER

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2509.04063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.04063 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:30:23.710832Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:51:28.966906Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T12:43:17.233166Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e864b4a-3641-41f4-a30f-edb30d1a5381 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.444500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.444500Z digest=sha256:ed1adac2fc5b52bb5a6b3b87ae3f13a168331e089c354435d9d35e1df6768f24

Observation 22c3cd8a-a31b-4e15-af7b-55d50703bf4e · outbound

This paper cites write newline.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.451025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.451025Z digest=sha256:806472c8eaecc4e9302cdbe278b4ff1388e3e97ce369b32f20a1f7c99f3c35ee

Observation 27ed1442-1a51-4744-899a-bc5d90be4341 · outbound

This paper cites B.; Jaakkola, T.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models B.; Jaakkola, T

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:30:24.520293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.457237Z digest=sha256:187c1da3ef5db25e4171c9ea5881222615606661c74dccbd0506906a63053927

Observation 57708748-ee58-4c4c-af3c-4baef8fcf1a6 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.463943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.463943Z digest=sha256:a519e13bd3fc3bdef80af2283a262b099fdc8e0c7799cdc9b4b6a2c3d5829d14

Observation 49033faf-97a7-4f53-9d63-22125454ec02 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.469543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.469543Z digest=sha256:42ce2ab13ba6117dce5750df4979a472ecabe696142502eb5ed1a3b6f722c5be

Observation aebd0cbe-21f0-4217-8ae9-d1bdccd21934 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.475841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.475841Z digest=sha256:0d70f2ea00306f15683f549c7b9e9a76a4563c1372706eb9887502f98935adb2

Observation fa1a897c-afaf-4e44-9a1c-d5a82059ee22 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.481852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.481852Z digest=sha256:fce0033367fe1f87a48048d57e6a7038eb20d8f35e78791035258b54fe407736

Observation d63c2495-1ddf-4d7f-88f7-592d5332d574 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.488324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.488324Z digest=sha256:81fe41fcb6dfba4be3766e9d638b7d0ea142d052dfa82269c3e20091f1815bbb

Observation c214714c-c744-48ab-87ce-77d35e271900 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.490508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.495241Z digest=sha256:7eca8473756f5bf2f5e8a3e690b2cdddd7edd9f685257e9a3803907842fdcd4e

Observation c308da86-c944-4c17-bc8f-eba9bc4d20ae · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.471335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.500979Z digest=sha256:ab034fb971ec8d52439f4eadadf18736785cec000ae7986c653eb982fb1127bc

Observation bb9929da-5aa4-48c0-abff-d109b0443be8 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.454442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.508419Z digest=sha256:28b94944cace17576a4b47444b3407e338dfbfd3bd430815b5b78c1afedbab81

Observation dfe423ff-3d87-449c-a8c7-b60679f9a6a8 · outbound

This paper cites Neural Ordinary Differential Equations.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Neural Ordinary Differential Equations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.515915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.515915Z digest=sha256:f145b4fc3d28764da95323ff66d8631dfd642547e0e5049fa9923fa255d6b79c

Observation f3e1c7e4-be1d-42dd-a51d-db9aa39d679a · outbound

This paper cites FDPP: Fine-tune Diffusion Policy with Human Preference.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models FDPP: Fine-tune Diffusion Policy with Human Preference

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.521640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.521640Z digest=sha256:98886a918a4ea3f1bceeffd19448320c571c89de9aabec735adb5f9255ecee64

Observation 7dace495-c2da-4461-814f-4faa88315a76 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.437323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.527844Z digest=sha256:143a150d6f5cfcabd4aa097fb9e62c537805712d1250e27eb460e09a889e5b8c

Observation f6f29c4d-4779-4f8d-9a3b-b6b65dca152f · outbound

This paper cites S.; Lynch, C.; Chowdhery, A.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; et al.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models S.; Lynch, C.; Chowdhery, A.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:30:24.419516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.533520Z digest=sha256:c8085e5370d94a967e397772d10f8f5035efe5a49672ef14b5c0e79779c4d003

Observation 35ba3196-ca77-4c1e-8e83-1549cb75e2ac · outbound

This paper cites Improving Vision-Language-Action Model with Online Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Improving Vision-Language-Action Model with Online Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.538701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.538701Z digest=sha256:afb6a214fd9a04aaf324f12ce0ba9b0b3d627c9568d206694a6d743131b40fb8

Observation 06c6519d-9fdc-4bb8-9e5e-546e23baea4f · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Denoising Diffusion Probabilistic Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.546728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.546728Z digest=sha256:9b745661af7e222c55303e8ed9f7d22a3bf6c50d55230390888316e57a70943f

Observation ce01b314-0f73-42dd-81e9-bb2039a2668e · outbound

This paper cites Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.552915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.552915Z digest=sha256:28e9f55d62ac59373ef4d481550a434947b38042224a54aa2287e2ff2851423f

Observation 4f31f06f-4b50-42b9-ad72-780caa8b1bbb · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.401852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.558652Z digest=sha256:89629c4554d12a8d5e4259a5a7b424ef6e4ad5de2d489cbd19ba1e13495af606

Observation a8d857f0-8d2b-45f4-a6b8-b9b67151b6f9 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models OpenVLA: An Open-Source Vision-Language-Action Model

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.563766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.563766Z digest=sha256:916ff61d2663e87b89425953be90c7aa786af29fd715af1993df12e1e4142131

Observation e284bcf8-56be-4f37-b6e2-3a5bd1de3fd5 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.382885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.570311Z digest=sha256:77069d6edcc82a30d1320752c685a201588f6334d1461797cc31600b4ea4f9e1

Observation 79743d51-d3a7-4fd5-a3a7-23ca43a66281 · outbound

This paper cites Evaluating Real-World Robot Manipulation Policies in Simulation.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Evaluating Real-World Robot Manipulation Policies in Simulation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.576111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.576111Z digest=sha256:ac4554479d3a7feee485938cb447b56fd042dd6f86ae65a11605198aec4ab126

Observation fe88a4da-8409-4ba1-9ecd-04a26d19ce36 · outbound

This paper cites Flow Matching for Generative Modeling.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Flow Matching for Generative Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.582107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.582107Z digest=sha256:a215ee94c19f7021fb2203ee800c70ad2e9b36917948c2cd872fe569008c1d3d

Observation 1bdee1ee-8647-4c50-93e4-15dd49ab6174 · outbound

This paper cites LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.587313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.587313Z digest=sha256:896c8b3a75f27fb1e047a1c6b93b4265d35ab9d39b5db8b5b7c4a8a98fa44a12

Observation 2a368fcf-5caa-4743-9d43-14b625537fab · outbound

This paper cites Decoupled Weight Decay Regularization.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Decoupled Weight Decay Regularization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.593021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.593021Z digest=sha256:3c06017cbe843f875048f074b060b63b5b870015f6e7477e7ffb9b462a7c4769

Observation 6e05dbf2-4666-4a20-870d-2f28538071a6 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.361360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.598619Z digest=sha256:d4b182043957ab2a5309a4daf34d042b72fb985356339c855dbd73a266acbfdb

Observation 7cc1bd6a-b221-4f54-b7e7-2c0543864e5a · outbound

This paper cites VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.605483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.605483Z digest=sha256:e9c2f7aab10bbf48554a4f230219bc2b4539fb17dc9fc8b14ebb9e05a61e80f2

Observation cdd8c86c-0b3b-4f1a-af6f-f9b4914be986 · outbound

This paper cites Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Policy Agnostic RL: Offline RL and Online RL Fine-Tuning of Any Class and Backbone

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.611023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.611023Z digest=sha256:b0a29af8e1f5aad6d663db0d092e97396b646fbe0956d742115468be77143fb4

Observation 2f15b711-1537-4c34-acf5-e1a2688baef3 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.343397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.616193Z digest=sha256:c895d92adfc7971b018eb9840ba125aa45a873607763ac2fba70945ed7aa4f08

Observation f81a2e7d-54e7-48a4-a8d8-e0ffe599cf5c · outbound

This paper cites Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.621321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.621321Z digest=sha256:bd84ab23a28619a6f2f9bd8f20ecd91a00fe3f7cfe2d389f91c04136e7b008d0

Observation 536ba1d1-401a-476c-bf6a-ebd0758db317 · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.326271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.626576Z digest=sha256:ce1a666c432388cddd266b061318b6b2bef082e5601bdc7084e3a795d7e8c02d

Observation ffee3e7b-3f88-4282-9ee9-d2b050e62ed1 · outbound

This paper cites Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.635769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.635769Z digest=sha256:fd49da79ad2a831e0996b187343661e3ca64e0e5eaa80d5bab61f33fee3bd9d9

Observation d44e2bf0-3e45-4016-bd0e-0320436923be · outbound

This paper cites Proximal Policy Optimization Algorithms.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.641166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.641166Z digest=sha256:8930ac92dbf4c337b124a6a713bb99fbc31f40bde1352b6d241953d18ab80032

Observation 8bc3f0bd-77b6-4e4c-9cbe-07689334638b · outbound

This paper cites Interactive Post-Training for Vision-Language-Action Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Interactive Post-Training for Vision-Language-Action Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.646643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.646643Z digest=sha256:ca22c73d777f3bc16d9525f2ae5a99069c5aa2393805c2dcb8778c752168aa8b

Observation 63d5175f-027b-465e-8d22-c3e5fbc4fa6b · outbound

This paper cites Octo: An Open-Source Generalist Robot Policy.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Octo: An Open-Source Generalist Robot Policy

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.652369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.652369Z digest=sha256:da6f8115f7d74e8a233fce75bd8cd2c47a4c70cb46d841b7cfd31302812c4670

Observation b8a9a791-599a-409c-89db-a0c000554697 · outbound

This paper cites J.; and Zhou, M.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models J.; and Zhou, M

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:30:24.308875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.658056Z digest=sha256:adca9576c3132a98eabafd1ca0a21577ab386a34e6579c142dff7775e74a2157

Observation dd6f62e4-7bcf-4126-aae7-ce8dc1fdf77f · outbound

This paper cites an unresolved cited work.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:30:24.289673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T10:30:23.664002Z digest=sha256:fe5218e546948fde2008469fc987371365b593787977bc2886d14c924d5137f9

Observation c66a75f6-4f55-462d-b786-087c9b8bb200 · outbound

This paper cites ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models ReinboT: Amplifying Robot Visual-Language Manipulation with Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.670740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.670740Z digest=sha256:4001d023298cc9d2c9c21616e61787f78e2ad24cb6e0060acd3175c372e5dbbc

Observation 94d2adb0-a693-488f-9d59-14070dd5b872 · outbound

This paper cites Energy-Weighted Flow Matching for Offline Reinforcement Learning.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Energy-Weighted Flow Matching for Offline Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.678531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.678531Z digest=sha256:6e41cb4dcfc5eae08eed05d52e2f7f29b4a3ca37f7caa4f83e1a1bb755555f4b

Observation 02898c68-a010-4aa3-adf2-0cfe4243bf80 · outbound

This paper cites GRAPE: Generalizing Robot Policy via Preference Alignment.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.688049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.688049Z digest=sha256:04abe58b795eba90834a50f62c5983326732e61116af53f4ddc38d412639a0cd

Observation 8630cad3-afcf-41be-86f7-535f15cd85e6 · outbound

This paper cites MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.695864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.695864Z digest=sha256:8412c0cdb29e4bb0c17e125861db3415c72ae412f986afdb5528022965da9f28

Observation 6ef7b8c5-a478-446f-a696-a7c7126604be · outbound

This paper cites Guided Flows for Generative Modeling and Decision Making.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Guided Flows for Generative Modeling and Decision Making

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.703140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.703140Z digest=sha256:8aab147c1da5222eb17bcbf88d0444fe8d4e5fd6012b78d4248a8ba20b861c6b

Observation 5f98f877-341c-4bed-be13-e0f6251079c8 · outbound

This paper cites Reinformer: Max-Return Sequence Modeling for Offline RL.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models Reinformer: Max-Return Sequence Modeling for Offline RL

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.710832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.710832Z digest=sha256:beb2844c2e86bbad20bd72d11854d5901745b99ac85fc0a2161e42a83feed6bc

Pith citing papers

Observation 0515b405-0150-47b2-86a7-06e4b120a5ca · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T07:31:02.908493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:ce9eabc9990d8803f5ba3ab592935d9c7dc54cfb0921b056928bc037d2a10240

Observation 523fe429-f1b8-4cb6-8c19-09b961db3bf4 · inbound

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing cites this paper.

HapticVLA: Contact-Rich Manipulation via Vision-Language-Action Model without Inference-Time Tactile Sensing Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T05:51:28.966906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:51:28.966906Z digest=sha256:0d58a646c2af99d04b607b7d7871e550ae0ca11128a9ff9b79e5e5feb06da3d0

Observation 6ef27ad8-2026-4fda-934b-461fa5a9a422 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.235496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:b675b7adb65bd3bd460e899c171ea4b9d7b54d850baa3e310f51dece700746cf