Pith. sign in

Paper Citation Record · LEDGER

GRAPE: Generalizing Robot Policy via Preference Alignment

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2411.19309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19309 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:55:50.796899Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d92e50b6-da95-4b76-a598-0a96b768e662 · inbound

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment cites this paper.

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T19:55:50.796899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:55:50.796899Z digest=sha256:127ed654ac88aa3223f75a5c0a75c7d470467f7f0539567a32f1d70390fff069

Observation cb8b47de-5906-4be4-b685-a637e828abe3 · inbound

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation cites this paper.

MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:50:57.959258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:50:57.959258Z digest=sha256:156ab2e789e1978569c033d3633c887e41bfb88dd6e5f844d4484a4de8fc9d84

Observation 83a3fcea-f4b4-4d05-886d-741da72770f3 · inbound

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control cites this paper.

DexVLA: Vision-Language Model with Plug-In Diffusion Expert for General Robot Control GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:48:48.806390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-14T19:48:48.725800Z digest=sha256:da98b845047ee40f3282faf2b30f2ed9d5481f3ec58da54eefe2b0beddad0776

Observation d799988c-429d-4e5f-9812-7004b9454f9e · inbound

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning cites this paper.

SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:32:22.363054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T01:27:33.123243Z digest=sha256:1c7b6871fa77135e3d5ecf3a2e6719d612d298940aa0fbc43983feac8a140aef

Observation 2fe91573-9106-4706-b7c4-aa0c4edc2f88 · inbound

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems cites this paper.

AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:09:24.594651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T15:09:24.367362Z digest=sha256:a8b3dd98f4440295eaddf1eb492706ec1e0def5aa735242ae7a0ba0a696ebc47

Observation a6898c71-0575-4ffd-8d6e-122c1bd2b242 · inbound

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions cites this paper.

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:28:06.959009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T15:28:06.883492Z digest=sha256:7cf145ed0f12e99d29609fd56db08b6aed4c0834c7dc813b405e07cf430bacea

Observation a50457bd-f4cd-4f9a-a775-a7e4dfb9c6c6 · inbound

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning cites this paper.

VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:55:40.398126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T12:55:40.245908Z digest=sha256:2dc586d5de68399b297ab74eaa9e3fdcffa834de3ac368b30b181965ec6970e7

Observation 4bface85-924a-42ee-a04b-36476e85f515 · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.572064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.572064Z digest=sha256:1ebb12fa93977711a8ad0f230c165d7803c10f6a65a442f587a15fc20fcb5c96

Observation 29e67213-daad-427c-b70d-6acbf6f1b574 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 258

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:37.279568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:37.279568Z digest=sha256:26a6fa75107668d02b8bb6a7fbe2501155a17843b29d6f41c83d717c5105b94b

Observation 8878e613-9cc9-4eda-923c-404de6112c0d · inbound

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey cites this paper.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 177

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.109526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:28:15.818016Z digest=sha256:6880e09e0233c40034f4591b77b1fccaecb2721d4b7adae13c519d9d2ed3ee9f

Observation 02898c68-a010-4aa3-adf2-0cfe4243bf80 · inbound

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models cites this paper.

Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:23.688049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:23.688049Z digest=sha256:04abe58b795eba90834a50f62c5983326732e61116af53f4ddc38d412639a0cd

Observation fd5ff45e-a0ba-4ba4-952a-a685a8f4638b · inbound

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning cites this paper.

SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:02:11.464210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T08:02:11.189795Z digest=sha256:67b9faeeaeb9aa37e35e9733a34ed4497cb0ca2db0696d7b3898638d1c33e0ff

Observation 54f664df-f197-40ff-89df-f5497b457f57 · inbound

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models cites this paper.

Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:24:57.958256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:24:57.958256Z digest=sha256:5d30e5d57b119e8fa324d8560a9dcb3ef4c43275a925e76dffbda55cb5780846

Observation 40ef77c0-734c-44db-9fd0-8c8f4fb483ea · inbound

Reflection-Based Task Adaptation for Self-Improving VLA cites this paper.

Reflection-Based Task Adaptation for Self-Improving VLA GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:31:02.933528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T07:28:11.187479Z digest=sha256:aaedb45fdd39d96af6c2cffa2c4bac1cdf535300c5469c7b10b8348a7bc0be63

Observation a64082d7-3ec8-4d36-bd0f-2c30aa744d6b · inbound

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation cites this paper.

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:10:57.868145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T06:10:47.309028Z digest=sha256:0d9322fcc406942993f31ec037f077da8229a6dd149aaac25b3a3dad153812fd

Observation 1dd654af-0a47-4061-a394-7a1a3f3bf613 · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.387695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:8d83f9189a3c5b863c0445ba7fc95250e1744448e4e9b384825296a0c2a914fc

Observation eb701436-3068-4bac-a2b7-ca52cf2f6edc · inbound

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models cites this paper.

DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:10:48.869856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T03:09:09.713822Z digest=sha256:0bc2ca4e21b518ab53d0561ae0e627231e1934f918919e3b05426bf83de49090

Observation fd542b69-2fa3-45a5-a2df-22e01d8028ab · inbound

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models cites this paper.

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T13:53:27.515205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:53:27.515205Z digest=sha256:cea15b36626b2a3e2aa9ff1d2d3bf906a6e6b88885ad20f13b259e2aaa81478d

Observation 62cb33a6-30c1-48ce-b3d1-b892d91f022f · inbound

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation cites this paper.

TwinRL: Digital Twin-Driven Reinforcement Learning for Real-World Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.900849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T13:13:53.818915Z digest=sha256:779e3d60f54401c50933eb3962a49b2b3facd808592725e2270d62fe9bf0317b

Observation 904c8a16-d0a8-4b19-8bae-3e0284a62a0e · inbound

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning cites this paper.

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T16:10:12.689957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:10:12.689957Z digest=sha256:c9835f8c812756717d8b25dfacca2b2af54705e5f005522587e04f88403cc776

Observation 42866686-c264-452a-a320-63705cc5313b · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:30.009180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T05:47:17.494531Z digest=sha256:e393dcb58179b0cf50ead952aefffd3205bc3b4bb2aa0adbad531ce71d436af8

Observation e7e21ad5-2002-4d86-9cfb-3157990b3a32 · inbound

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning cites this paper.

LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:17:03.135989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T03:00:26.352130Z digest=sha256:3dccc167e5bd7478dccb1b096b706431eea6b3c9c1b5e7c68857908288413e08

Observation 5605a0c1-a662-42a1-acba-56c0b3227373 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:36:10.861354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T19:31:38.069592Z digest=sha256:4b2a40eafa1109bd50151cd6bbf36e7b520a9a5b9380181765f96c2c99428dbf

Observation 8446b9d6-3128-4dda-b251-182de97f9328 · inbound

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies cites this paper.

Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:15:32.297095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T08:05:47.128354Z digest=sha256:fb8a53c3042712a3f5e897446e9f9782818dfd8cba8ec8406557b9d0cd4d6385

Observation 14c156f9-c290-4a3b-b2ff-a42e8914e864 · inbound

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation cites this paper.

ProcVLM: Learning Procedure-Grounded Progress Rewards for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:41:19.940732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:40:22.192349Z digest=sha256:537177f2fe504f6e59400863ed9d7a8859dd597a139f88a9d14e433d34c402fa

Observation 406737e0-8593-4ee5-98f8-7b550cf92ba6 · inbound

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models cites this paper.

RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.984235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T04:39:32.715650Z digest=sha256:9726244dca11374cd2b3bb7982865d7f7c192a1b1341977b81c2e510428c253b

Observation 5e16325d-038d-4ed9-998a-680fd90a8f0c · inbound

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation cites this paper.

From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.884932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T04:44:03.661688Z digest=sha256:31c7ced5aea6f606df23a7992652040e3410728451bdad3eead4e7943b3f3bba

Observation c38a1916-b160-4388-893a-8d8bf55445c7 · inbound

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization cites this paper.

DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 191

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:17.293277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T12:39:50.004269Z digest=sha256:83bcd8def76acb8ae2198d5adfe0ce09c2da275d1d5d8a1f9c1182e0cbd099f4

Observation 5114bd6a-3171-4846-b807-df20d8a2dfa4 · inbound

PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models cites this paper.

PAPO-VLA: Planning-Aware Policy Optimization for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.383440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T05:12:26.907425Z digest=sha256:9772177a4e3e7b45c885d46eaf9fa8d46427eca6c559338b294ac500bc103c90

Observation 9a8982b7-5c89-4284-bcae-25bde5ea9884 · inbound

Position: Good Embodied Reward Models Need Bad Behavior Data cites this paper.

Position: Good Embodied Reward Models Need Bad Behavior Data GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.973243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T17:18:17.337336Z digest=sha256:566af50fdad1c00fc046d201132daf7a5264886ecfec7cc3bd6f6180129e471d

Observation 02a91f91-65be-4ab1-a6f7-ca0aaa2d4ece · inbound

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization cites this paper.

FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T09:06:49.389175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T05:38:11.089753Z digest=sha256:58133445ad5b742cf5c14b8f597487324707cb6ed4518d08db5a95e979ee51ec

Observation 7f8b034d-70c4-4bbb-b528-7fe26894c452 · inbound

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning cites this paper.

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.728427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:46:59.746745Z digest=sha256:a1f0aee441120359fe103f2af6742b7797eac654f40630d71c682a98b1bead67

Observation 627851db-f62c-4f08-9b3e-9bb97f8af871 · inbound

Foresight: Iterative Reasoning About Clues that Matter for Navigation cites this paper.

Foresight: Iterative Reasoning About Clues that Matter for Navigation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.175007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:35:05.441401Z digest=sha256:0136cf4daea8356dcb6ac7cc7627248f0732c32e0ada978d23adef67585a6a6d

Observation f7277805-4317-4e28-9e61-fc005ca76c40 · inbound

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack cites this paper.

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T11:31:37.690666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:31:37.690666Z digest=sha256:e78fd1bcd2ba963e03840dd1ac2b8b9b5523e4b2c0b26dadba4769edcd4bf02f

Observation 62cfe2bf-a65b-4d10-9fe0-c2e626669ab7 · inbound

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model cites this paper.

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:44.353030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T04:13:22.598591Z digest=sha256:f030185ebbccad4e58c18e3cc56356af5b84610389b83e5ac447d4dfff76513d

Observation dd067f52-1652-4806-8e97-4d554a3ac16d · inbound

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model cites this paper.

SafeDojo: Safe Reinforcement Learning for VLA via Interactive World Model GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T04:20:31.958581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T04:13:22.598591Z digest=sha256:9c8bb61dd12b0c52cdfa9b4bf308a414b75cd4c219722d2ae6665545e78e29f9

Observation 3e82a913-f3bb-4323-ad4c-b10b8512e9dc · inbound

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning cites this paper.

HiL-ResRL: A Model-Agnostic Finetuning Adapter via Human-in-the-loop Residual Reinforcement Learning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.655525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T08:45:04.483217Z digest=sha256:86826dd6b48cae597915ff9c456b35e21a6b8fe33e639a6654d203aca2bb1f98

Observation 37132ba4-8cb9-454d-87ad-233ba7680559 · inbound

Learning Process Rewards via Success Visitation Matching for Efficient RL cites this paper.

Learning Process Rewards via Success Visitation Matching for Efficient RL GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.474130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:20:35.062060Z digest=sha256:3797fd0bfe53a43a05157adf5a3ce5f8da444a243e6e61ac605da416bbff5891

Observation 276ca8c9-6be5-404b-8e94-6c71abf6c499 · inbound

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models cites this paper.

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.062167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-25T20:22:16.508280Z digest=sha256:f861fbf48f196407bc9263defdc8f97cb0bf3fd609163bf233f5531a82243719

Observation f7e59c4a-1b10-4957-ab3e-fd326e41c4b5 · inbound

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models cites this paper.

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.163124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T06:02:15.781538Z digest=sha256:99ecd0a3ee0cd70150153cbee46cdf16c7c5f4dfb7d43460e5f6c54d23a268fb

Observation 5c19d9a5-74d1-4f82-8f75-e53ecebe2fc2 · inbound

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning cites this paper.

Rethinking Foundation Model Collaboration: Enhancing Specialized Models through Proxy Task Reasoning GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:45:29.789741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T06:38:44.170473Z digest=sha256:e6f2462eb557a432635aef383ecbeceb00db056e68bbd9cc445899de618058ae

Observation 534908e9-3bd8-4b9e-a92e-584e4d8aa8a2 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.966122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T04:58:28.971536Z digest=sha256:130ac2d013e9815bc83c939686946ca195e3e71fc864c0a525f1b9a13e9cd5ab

Observation 2d48f57a-1bc6-4b95-bcde-dd18adfd100a · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:18.028851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:18.028851Z digest=sha256:55b9bbf87e73a24171d219ec036c24b59d2fb129422080b6decef1395dfdd7d4

Observation 1b5e5b47-3b10-4bcf-8972-443f065ab236 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:19.142604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:32:19.142604Z digest=sha256:994e0c7dbc323bc192e5fbb97e018d625176822b99954ad26f386b3066fc17d3

Observation 2f0f7346-005b-42ac-ba88-c2df11dc3a3b · inbound

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models cites this paper.

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T14:45:01.915411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T14:45:01.915411Z digest=sha256:09dfcdf28b6f37540e762a8c2bec2cbca38f2b6fd36c8660f18372c959fbf3f7

Observation 8247ecf9-282f-4cea-99d9-e682e690fe3f · inbound

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models cites this paper.

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:23:14.040221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:23:14.040221Z digest=sha256:7e0de2bfb72e98f0a2e18c46fe886b82d1eb9c69efa94f8f87a4705fc184478f

Observation a95f753a-b662-479e-bbb8-80748f72cebb · inbound

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy cites this paper.

RedFlow: Redirect Failure into Action-Level Corrections for Flow-matching VLA Policy GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:32:25.304633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:32:25.304633Z digest=sha256:0963fd6460216db4ae10739e908b1b6b1d2466681d0aea7dd89641fe488d8639

Observation 97963ac0-26fe-4a7e-b568-b59535017128 · inbound

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation cites this paper.

Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation GRAPE: Generalizing Robot Policy via Preference Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T19:38:41.768698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:38:41.768698Z digest=sha256:a54f478f897fc98e15da0af5f8510dc120fb59df030a61152593b87f4e4618e7