Pith. sign in

REVIEW 2 major objections 5 minor 26 references

Effective robot motion skills come from predicting how 3D geometry evolves, not from matching pixel patterns.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 14:42 UTC pith:427BSKMI

load-bearing objection Solid empirical recipe: VQ motion codes trained by single-view pointmap future prediction beat strong 3D baselines; the 'true physical 3D' claim is a bit oversold but the numbers hold. the 2 major comments →

arxiv 2607.04714 v1 pith:427BSKMI submitted 2026-07-06 cs.RO cs.AI

Geometry-Aware Motion Latents for Learning Robust Manipulation Policies

classification cs.RO cs.AI
keywords robot manipulationmotion latentspoint cloudsgeometry-aware learningdiffusion policy4D dynamicsvector quantizationsingle-view RGB-D
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that reusable motion patterns for robot manipulation should be learned by forecasting how a three-dimensional point cloud changes during an action, not by reconstructing the next camera image. Discrete latent codes trained on that four-dimensional geometric objective are forced to encode physical transformations rather than appearance. The resulting policy, GeoMoLa, reaches state-of-the-art success on standard benchmarks from only a single RGB-D view and improves few-demonstration real-world control in clutter. Ablations show that removing geometric prediction hurts performance far more than removing visual prediction, and the same codes produce consistent motions when applied to new scenes. The central claim is that motion latents for control emerge more reliably from three-dimensional effects over time than from pixel-level patterns.

Core claim

Effective motion latents for robot control emerge more reliably when the learning objective is to predict future three-dimensional point-cloud geometry than when it is to reconstruct visual appearance. This geometry-first objective produces discrete codes that capture physical transformations, yield state-of-the-art single-view success on manipulation benchmarks, and remain consistent when transferred to novel scenes.

What carries the argument

Geometry-aware motion latents (GeoMoLa): discrete vector-quantized codes trained by conditional diffusion to predict future pointmaps from current RGB-D and language, then used to condition a 3D denoising transformer that outputs 6-DoF action chunks.

Load-bearing premise

That forecasting single-view pointmaps is enough to force the discrete codes to capture true physical 3D motion rather than camera-specific depth cues or leftover appearance.

What would settle it

If a policy trained only on RGB future prediction matched or beat full GeoMoLa success rates on rotation-heavy RLBench tasks and real cluttered stacking, the claim that geometric prediction is the key driver would be falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Single-view RGB-D policies can match or beat multi-view reconstruction methods when motion latents are trained on geometric evolution.
  • Ablating geometric prediction degrades success far more than ablating RGB prediction, so spatial dynamics are the primary signal for manipulation skills.
  • The same discrete codes produce consistent motion types across different visual scenes.
  • Few-demonstration real-world policies improve most on cluttered and occluded tasks that require spatial reasoning.
  • Motion latent learning should treat actions as continuous 4D geometric processes rather than 2D video patterns.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same 4D prediction objective could extend to deformable objects if surface or mesh representations replace rigid pointmaps.
  • Hierarchical planners could compose these geometric primitives into longer skills without relearning low-level dynamics.
  • Multi-view consistency checks on the codes would further test whether they are truly view-invariant 3D transformations rather than single-camera artifacts.
  • Large unlabeled robot video sets could be mined for motion primitives by back-projecting single-view depth instead of pure visual reconstruction.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces GeoMoLa, a two-stage framework that learns discrete motion latent codes (via VQ on a VLM encoding of single-view RGB-D + language) by training them to condition a diffusion model that predicts future pointmaps (and jointly RGB) rather than reconstructing current observations. These geometry-aware latents then condition a 3D denoising transformer policy that generates 6-DoF action chunks. The central claim is that the 4D geometric prediction objective forces the codes to encode physical 3D transformations, yielding SOTA single-view success on RLBench (84.7% avg over 10 tasks) and CALVIN ABC→D long-horizon chaining (avg length 3.60), superior few-demo real-robot performance on ALOHA (esp. cluttered/occluded tasks), and transferable motion primitives (qualitative cross-scene consistency in Fig. 3). Ablations (Tab. 4) show pointmap prediction drives most gains while RGB contributes little.

Significance. If the results hold, the work provides a practical and well-supported advance for motion-latent learning in manipulation: it shows that a future-pointmap diffusion objective produces more effective discrete codes than 2D video or static 3D baselines, delivers clear single-view SOTA numbers against strong multi-view and diffusion competitors, and includes a clean modality ablation plus real-robot validation with only 20 demos per task. The explicit credit for geometric vs. appearance modeling and the demonstration that the same codes induce consistent motions across scenes are useful contributions to the representation-learning side of robot learning. The approach is immediately usable (single RGB-D + language) and the empirical package (RLBench 5 seeds, CALVIN zero-shot, real ALOHA) is stronger than many concurrent latent-action papers.

major comments (2)
  1. [§3.2.2, Tab. 4, Fig. 3] §3.2.2 (Eq. 2 / L_pm_diff) and the interpretation in the abstract/§1/§5: the claim that the future-pointmap objective 'forces latent representations to encode actual physical motion rather than appearance patterns' (and produces 'physically consistent transformations regardless of visual context') is only partially supported. Pointmaps are single-view back-projections; the large drop when removing the pointmap branch (Tab. 4) and the small drop when removing RGB show that geometric prediction helps, but do not isolate view-invariant rigid-body motion from monocular depth biases, occlusion patterns, or residual appearance that co-vary with successful actions. Fig. 3 is purely qualitative (visual consistency of 'down'/'rotate'). A quantitative check—e.g., multi-view consistency of induced 3D trajectories, Chamfer distance of transferred pointmap predictions, or rigid-motion residual after
  2. [§4.1–4.2, Tabs. 1–2] §4.1 / Tab. 1 and §4.2 / Tab. 2: several strong baselines (GNFactor, ManiGaussian) are trained with 19 extra views while GeoMoLa and the main competitors use only front-view RGB-D at inference. The paper correctly notes this, yet the SOTA claim would be more robust if an ablation or re-run of the multi-view methods under the identical single-view constraint were provided, or if the gap attributable purely to the motion-latent objective (vs. the 3D denoising transformer backbone shared with 3D Diffuser Actor) were isolated more cleanly. The current numbers are still impressive but leave open how much of the 7–10 point lift is the 4D objective versus architectural parity.
minor comments (5)
  1. [abstract, §1, Fig. 2] Throughout (esp. abstract, §1, Fig. 2 caption): repeated missing spaces after method names ('GeoMoLaachieves', 'GeoMoLaframework', 'GeoMoLashows') and occasional capitalization glitches ('We now describe'). Clean for camera-ready.
  2. [§3.2–3.3] §3.2.1 / Eq. (1) and §3.3.1: the number of discrete codes ns, codebook size K, and patch size are given only in the appendix table; a short statement of the chosen values (and sensitivity) in the main text would help reproducibility.
  3. [Fig. 3, §4.1] Fig. 3 and Fig. 4: the latent codes are shown as integer tuples but never linked back to the codebook visualization or nearest-neighbor retrieval; a small quantitative transfer metric (e.g., success rate when swapping codes across tasks) would strengthen the interpretability claim without new experiments.
  4. [§4.4, Tab. 3] §4.4 / Tab. 3: real-world results are reported over 10 trials; adding standard error or a note on variance would match the 5-seed reporting used for RLBench.
  5. [§2] Related Work §2: LAPA and Moto are cited for 2D motion latents; a one-sentence contrast on why their video-only objectives are insufficient for the geometric claim would tighten the positioning.

Circularity Check

0 steps flagged

No circularity: standard self-supervised latent learning + conditional policy, evaluated on external task success; no equation or claim reduces to its own inputs by construction.

full rationale

The paper's derivation chain is a conventional two-stage ML pipeline. Discrete motion latents z_t are obtained by VQ of VLM features and trained via a conditional diffusion objective that denoises future pointmaps (and optionally RGB) given history and z_t (Eq. 2 and L_pm_diff / L_rgb_diff / L_vq in §3.2.2). Those frozen latents then condition a separate 3D denoising transformer that predicts action noise (L_θ in §3.3). All reported numbers (RLBench success rates, CALVIN chain lengths, real-world success rates, Tab. 1–4) are external task-completion metrics measured after training; none is a re-statement of a fitted free parameter. Ablations simply drop one prediction branch and re-measure the same external metrics; Fig. 3 is qualitative transfer of codes, not a quantitative prediction forced by construction. Self-citations (GENIE, LAPA, 3D Diffuser Actor, etc.) are ordinary prior art and do not supply a uniqueness theorem or ansatz that forces the present results. No self-definitional loop, fitted-input-as-prediction, or renaming of a known identity appears. The pipeline is therefore self-contained against external benchmarks.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 1 invented entities

Empirical robotics paper. Load-bearing content is the modeling choice that future pointmap prediction yields geometry-aware discrete codes, plus standard VQ/diffusion machinery and the usual rigid-manipulation, language-conditioned, single-view RGB-D assumptions. Free parameters are ordinary ML hyper-parameters; no new physical constants. Invented entity is the GeoMoLa motion-latent construction itself.

free parameters (6)
  • codebook size K and code dimension = K=64, dim=32
    Set to 64 codes of dim 32 (Tab. 7); controls capacity and discreteness of motion primitives; chosen by authors, not derived.
  • number of discrete codes per observation ns / patch size = patch size 16
    Patch size 16 and resulting sequence length determine how motion is tokenized; architectural free choice.
  • VQ commitment coefficient β and NSVQ noise schedule
    Standard VQ-VAE training knobs that affect codebook utilization; not fixed by theory.
  • action-loss weights λ_p, λ_r, λ_g = tuned
    Component weights for position L1, rotation L1, gripper BCE; tuned (§3.3.2).
  • diffusion steps / noise schedule for pointmap, RGB, and action models = 100 / 25 steps
    100 (RLBench) / 25 (CALVIN) action diffusion steps and corresponding α schedules are design choices that affect reported success.
  • observation window w and action horizon h
    Temporal context and prediction length for both future-prediction and policy; chosen for each benchmark.
axioms (5)
  • domain assumption Single-view RGB-D back-projection yields a pointmap that is an adequate geometric state for learning transferable manipulation motions.
    Invoked throughout §3.2; success of the method rests on this representation being rich enough without multi-view fusion.
  • ad hoc to paper Discrete VQ codes trained to predict future geometry will cluster into reusable, semantically consistent motion primitives.
    Stated motivation in §3.2; supported empirically by Fig. 3 transfer but not proved.
  • domain assumption Manipulation tasks of interest are dominated by rigid-body geometric transformations (paper notes focus on rigid-body).
    Conclusion and method scope; deformable cases left to future work.
  • standard math Standard DDPM / VQ-VAE / CLIP-ResNet / Mini-GPT components behave as in prior literature and can be composed without pathological interference.
    Background machinery cited from Ho et al., van den Oord et al., Radford et al., Zhu et al., Ke et al.
  • domain assumption Language-annotated demonstration distributions (RLBench, CALVIN play data, 20 real demos) are representative enough for the reported generalization claims.
    Evaluation protocol §4; zero-shot env D and real-world layout shifts assume this.
invented entities (1)
  • Geometry-aware discrete motion latents (GeoMoLa codes) no independent evidence
    purpose: Abstract reusable 3D motion primitives that condition a diffusion policy and are trained by future pointmap prediction.
    Core proposed representation; independent evidence is only the paper’s own transfer visualizations and task metrics, not an external physical measurement.

pith-pipeline@v1.1.0-grok45 · 24151 in / 3529 out tokens · 31773 ms · 2026-07-11T14:42:12.587903+00:00 · methodology

0 comments
read the original abstract

Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action abstractions require understanding three-dimensional geometric transformations. Here, we introduce GeoMoLa (Geometry-Aware Motion Latents), which learns discrete motion latent codes by predicting how point clouds evolve during manipulation rather than reconstructing visual observations. This four-dimensional objective -- spatial geometry changing through time -- forces latent representations to encode actual physical motion rather than appearance patterns. GeoMoLa achieves state-of-the-art performance using only single-view RGB-D input, while existing methods require multi-view reconstruction, succeeding across diverse manipulation benchmarks. Our ablations reveal that geometric prediction is the key to driving performance, quantitatively validating that manipulation depends on spatial understanding. Furthermore, the learned codes exhibit effective motion abstraction: applying them to novel scenes produces physically consistent transformations regardless of visual context. Our real-world experiments also confirm this robustness capability, achieving robust manipulation with minimal demonstrations in cluttered environments where geometric reasoning determines success. Thus, we demonstrate that effective motion latents for robot control can better emerge from understanding motion through its three-dimensional effects rather than pixel-level patterns.

Figures

Figures reproduced from arXiv: 2607.04714 by Leonidas Guibas, Ming Hu, Ruizhe Liu, Yanchao Yang, Yijia Weng, Yunchao Zhang.

Figure 1
Figure 1. Figure 1: Policies trained with 2D and 3D motion latents. When encountering novel and cluttered scenes (left) that differ from the training distribution, the policy trained with 3D motion latents (top) demonstrates superior task performance with more robust control, e.g., enhanced reaching and grasping accuracy. during manipulation. By training latent representations to forecast future geometric states rather than r… view at source ↗
Figure 2
Figure 2. Figure 2: GeoMoLa framework. (a) Geometry-Aware Motion Latent Learning: RGB-D observations and language instructions are encoded into discrete motion latents via VQ-VAE, trained by predicting future pointmaps and RGB images. This self-supervised objective ensures latent codes capture 4D dynamics (3D geometry over time). (b) Latent-Conditioned Action Prediction: The previously trained motion latent encoder and the co… view at source ↗
Figure 3
Figure 3. Figure 3: Cross-scenario motion latent consistency. Visualization of predicted future observations when conditioning the same latent code on different initial scenes. Left: Latent code [1, 1, 5, 7] consistently produces downward motion. Right: Latent code [0, 3, 2, 2] consistently generates rotational motion. available play data including unannotated sequences. (iii) Closed-loop methods: Clover (Bu et al., 2024) enc… view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative comparison of future observation prediction. ManiGaussian requires 19 additional training views and produces blurred predictions with geometric inconsistencies (see distorted object boundaries). Env A Env B Env C Env D Training Test “move the switch to turn on the light bulb” “push the button” “place the red block in the slider” “push the blue block to the left” [PITH_FULL_IMAGE:figures/full_f… view at source ↗
Figure 5
Figure 5. Figure 5: llustration of the four different environments in CALVIN (Mees et al., 2021). in both RGB and depth modalities. The generated frames faithfully capture the geometric displacement of the drawer and the manipulator trajectory over multiple timesteps, showing smooth and physically plausible motion progression. Importantly, the predicted depth maps remain well aligned with the RGB predictions, indicating that … view at source ↗
Figure 6
Figure 6. Figure 6: Future observation prediction on Calvin (Mees et al., 2021). … … … … … … Stack cups Place dish Stack cubes Put on shelf Place cube Clean cups [PITH_FULL_IMAGE:figures/full_fig_p015_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Visualization of the real-world experiment setup. 15 [PITH_FULL_IMAGE:figures/full_fig_p015_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: More qualitative results on real-world experiments. The pointmap VAE is initialized by the RGB VAE. Before the latent action learning, we first finetune it on the video sequences of demonstration for a few epochs. Latent Action-Conditioned Future Prediction. We generate future latents with a conditional diffusion model adapted from DDPM (Ho et al., 2020) with motion latent z t : hˆt+1:t+w pm = ψ diff h t−w… view at source ↗
Figure 9
Figure 9. Figure 9: Qualitative comparison in the real-world cluttered scene. At each diffusion step k, we add Gaussian noise: h pm k = √ αk · h t+1:t+w pm + √ 1 − αk · ϵ, ϵ ∼ N (0, I), and minimize the denoising objective: L pm diff = Ek,ϵ ∥ϵ − ϵϕ(h pm k , k, h t−w+1:t pm , z t )∥ 2 [PITH_FULL_IMAGE:figures/full_fig_p018_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

26 extracted references · 18 linked inside Pith

  1. [1]

    org/CorpusID:264172455

    URL https://api.semanticscholar. org/CorpusID:264172455. Blattmann, A., Dockhorn, T., Kulal, S., Mendele- vitch, D., Kilian, M., and Lorenz, D. Sta- ble video diffusion: Scaling latent video diffusion models to large datasets.ArXiv, abs/2311.15127,

  2. [2]

    org/CorpusID:265312551

    URL https://api.semanticscholar. org/CorpusID:265312551. Brohan, A., Brown, N., Carbajal, J., Chebotar, Y ., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Her- zog, A., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jack- son, T., Jesmonth, S., Joshi, N. J., Julian, R. C., Kalash- nikov, D., Kuang, Y ., Leal, I., Lee, K.-H., Levine, S., Lu, Y ., Malla...

  3. [3]

    org/CorpusID:254591260

    URL https://api.semanticscholar. org/CorpusID:254591260. Brohan, A., Brown, N., Carbajal, J., Chebotar, Y ., Choro- manski, K., Ding, T., Driess, D., Dubey, K. A., Finn, C., Florence, P. R., Fu, C., Arenas, M. G., Gopalakr- ishnan, K., Han, K., Hausman, K., Herzog, A., Hsu, J., Ichter, B., Irpan, A., Joshi, N. J., Julian, R. C., Kalash- nikov, D., Kuang, ...

  4. [4]

    org/CorpusID:260293142

    URL https://api.semanticscholar. org/CorpusID:260293142. Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y ., Hughes, E., Lai, M., Mavalankar, A., Steiger- wald, R., Apps, C., et al. Genie: Generative interactive environments. InForty-first International Conference on Machine Learning, 2024. Bu, Q., Zeng, J., Chen, L., Yang, Y ., Zhou, G., ...

  5. [5]

    org/CorpusID:272653959

    URL https://api.semanticscholar. org/CorpusID:272653959. Chen, Y ., Ge, Y ., Tang, W., Li, Y ., Ge, Y ., Ding, M., Shan, Y ., and Liu, X. Moto: Latent motion token as the bridging language for learning robot manipulation from videos

  6. [6]

    org/CorpusID:277151378

    URL https://api.semanticscholar. org/CorpusID:277151378. 9 GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation Policies Coleman, D., Sucan, I. A., Chitta, S., and Correll, N. Reducing the barrier to entry of complex robotic soft- ware: a moveit! case study.ArXiv, abs/1404.3785,

  7. [7]

    org/CorpusID:13939653

    URL https://api.semanticscholar. org/CorpusID:13939653. Edwards, A., Sahni, H., Schroecker, Y ., and Isbell, C. Imi- tating latent policies from observation. InInternational conference on machine learning, pp. 1755–1763. PMLR, 2019. Fu, Z., Zhao, T., and Finn, C. Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation.Ar...

  8. [8]

    org/CorpusID:266755740

    URL https://api.semanticscholar. org/CorpusID:266755740. Gervet, T., Xian, Z., Gkanatsios, N., and Fragkiadaki, K. Act3d: 3d feature field transformers for multi-task robotic manipulation. InConference on Robot Learning,

  9. [9]

    org/CorpusID:259308821

    URL https://api.semanticscholar. org/CorpusID:259308821. Goyal, A., Xu, J., Guo, Y ., Blukis, V ., Chao, Y .- W., and Fox, D. Rvt: Robotic view transformer for 3d object manipulation.ArXiv, abs/2306.14896,

  10. [10]

    org/CorpusID:259262273

    URL https://api.semanticscholar. org/CorpusID:259262273. Goyal, A., Blukis, V ., Xu, J., Guo, Y ., Chao, Y .-W., and Fox, D. Rvt2: Learning precise manipulation from few demonstrations.RSS, 2024. Ho, J., Jain, A., and Abbeel, P. Denoising diffu- sion probabilistic models.ArXiv, abs/2006.11239,

  11. [11]

    org/CorpusID:219955663

    URL https://api.semanticscholar. org/CorpusID:219955663. James, S., Ma, Z., Arrojo, D. R., and Davison, A. J. Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters, 5:3019–3026,

  12. [12]

    org/CorpusID:202889132

    URL https://api.semanticscholar. org/CorpusID:202889132. James, S., Wada, K., Laidlow, T., and Davison, A. J. Coarse-to-fine q-attention: Efficient learn- ing for visual robotic manipulation via discretisa- tion.2022 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pp. 13729–13738,

  13. [13]

    org/CorpusID:235606348

    URL https://api.semanticscholar. org/CorpusID:235606348. Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C. Bc-z: Zero-shot task generalization with robotic imitation learn- ing.ArXiv, abs/2202.02005, 2022. URL https: //api.semanticscholar.org/CorpusID: 237257594. Ke, T.-W., Gkanatsios, N., and Fragkiadaki, K. 3...

  14. [14]

    org/CorpusID:266361979

    URL https://api.semanticscholar. org/CorpusID:266361979. Liu, H., Lee, L., Lee, K., and Abbeel, P. Instruction- following agents with jointly pre-trained vision-language models.ArXiv, abs/2210.13431, 2022. URL https: //api.semanticscholar.org/CorpusID: 253098249. Lu, G., Zhang, S., Wang, Z., Liu, C., Lu, J., and Tang, Y . Manigaussian: Dynamic gaussian sp...

  15. [15]

    org/CorpusID:244908821

    URL https://api.semanticscholar. org/CorpusID:244908821. Mildenhall, B., Srinivasan, P. P., Tancik, M., Barron, J. T., Ramamoorthi, R., and Ng, R. Nerf.Communi- cations of the ACM, 65:99 – 106, 2020. URL https: //api.semanticscholar.org/CorpusID: 213175590. Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Sing...

  16. [16]

    org/CorpusID:263626099

    URL https://api.semanticscholar. org/CorpusID:263626099. Parker-Holder, J., Ball, P., Bruce, J., Dasagi, V ., Holsheimer, K., Kaplanis, C., Moufarek, A., Scully, G., Shar, J., Shi, J., Spencer, S., Yung, J., Dennis, M., Kenjeyev, S., Long, S., Mnih, V ., Chan, H., Gazeau, M., Li, B., Pardo, F., Wang, L., Zhang, L., Besse, F., Harley, T., Mitenkova, A., Wa...

  17. [17]

    org/CorpusID:231591445

    URL https://api.semanticscholar. org/CorpusID:231591445. Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gim ´enez, M., Sulsky, Y ., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A. D., Heess, N. M. O., Chen, Y ., Hadsell, R., Vinyals, O., Bordbar, M., and de Fre- itas, N. A generalist agen...

  18. [18]

    org/CorpusID:248722148

    URL https://api.semanticscholar. org/CorpusID:248722148. Rohmer, E., Singh, S. P. N., and Freese, M. V- rep: A versatile and scalable robot simulation frame- work.2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, pp. 1321–1326,

  19. [19]

    org/CorpusID:960339

    URL https://api.semanticscholar. org/CorpusID:960339. Shridhar, M., Manuelli, L., and Fox, D. Perceiver-actor: A multi-task transformer for robotic manipula- tion.ArXiv, abs/2209.05451, 2022. URL https: //api.semanticscholar.org/CorpusID: 252199474. Su, J., Lu, Y ., Pan, S., Wen, B., and Liu, Y . Roformer: Enhanced transformer with rotary position embed- ...

  20. [20]

    org/CorpusID:266379116

    URL https://api.semanticscholar. org/CorpusID:266379116. Vali, M. H. and B¨ackstr¨om, T. Nsvq: Noise substitution in vector quantization for machine learning.IEEE Access, 10:13598–13610, 2022. van den Oord, A., Vinyals, O., and Kavukcuoglu, K. Neu- ral discrete representation learning. InNeural Informa- tion Processing Systems, 2017. URL https://api. sema...

  21. [21]

    org/CorpusID:266374724

    URL https://api.semanticscholar. org/CorpusID:266374724. Xian, Z., Gkanatsios, N., Gervet, T., Ke, T.-W., and Fragkiadaki, K. Chaineddiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipu- lation. In7th Annual Conference on Robot Learning,

  22. [22]

    Ye, S., Jang, J., Jeon, B., Joo, S

    URL https://openreview.net/forum? id=W0zgY2mBTA8. Ye, S., Jang, J., Jeon, B., Joo, S. J., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y .-W., Lin, B. Y ., Lid ´en, L., Lee, K., Gao, J., Zettle- moyer, L. S., Fox, D., and Seo, M. Latent ac- tion pretraining from videos.ArXiv, abs/2410.11758, 11 GeoMoLa: Geometry-Aware Motion Latents for Learning Robu...

  23. [23]

    org/CorpusID:273351190

    URL https://api.semanticscholar. org/CorpusID:273351190. Ze, Y ., Yan, G., Wu, Y .-H., Macaluso, A., Ge, Y ., Ye, J., Hansen, N., Li, L. E., and Wang, X. Gn- factor: Multi-task real robot learning with general- izable neural feature fields.ArXiv, abs/2308.16891,

  24. [24]

    org/CorpusID:261396262

    URL https://api.semanticscholar. org/CorpusID:261396262. Ze, Y ., Zhang, G., Zhang, K., Hu, C., Wang, M., and Xu, H. 3d diffusion policy.ArXiv, abs/2403.03954,

  25. [25]

    org/CorpusID:268253298

    URL https://api.semanticscholar. org/CorpusID:268253298. Zhou, Y ., Barnes, C., Lu, J., Yang, J., and Li, H. On the continuity of rotation representations in neural net- works.2019 IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), pp. 5738–5746,

  26. [26]

    close the — jar

    URL https://api.semanticscholar. org/CorpusID:56178817. Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M. Minigpt-4: Enhancing vision-language un- derstanding with advanced large language mod- els.ArXiv, abs/2304.10592, 2023. URL https: //api.semanticscholar.org/CorpusID: 258291930. 12 GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipu...