REVIEW 4 major objections 5 minor 117 references
Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read By rewriting each linear layer of a diffusion-based vision-language-action model as a differential update $y_t = y_{t-1} + W \Delta x_t$ and computing only the nonzero bit-level parts of $\Delta x_t$, Deltoris cuts arithmetic by up to…
desk verdict A credible co-design with real engineering, but the 34x speedup rests on an unverified premise: latent activations across control steps staying bit-sparse inside a stochastic diffusion chain. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the temporal-difference identity $y_t = y_{t-1} + W \Delta x_t$ (Eq. 2), which converts a full matrix-vector product into a sparse additive update. Deltoris represents $\Delta x_t$ in sign-magnitude form and executes only the active 1-bit events on 1D systolic bit-serial PE arrays with output-stationary dataflow; the speculative-inference stage then batches multiple speculative frames through the same weights to amortize DRAM traffic that the differential computation would otherwise inflate by 80%. The identity plus the batched-verification scheme together carry the claimed reductions in both compute and data movement.
What would settle it
Record the layer-wise activation tensors of a diffusion-based VLA running in closed loop at 50–200 Hz, with the stochastic re-noise term $\sigma_t n'_t$ active, and compute the fraction of active bits in $\Delta x_t$ for every linear layer. If that fraction is near 50% rather than the >90% needed for the claimed 92.9% operation reduction, the reported speedups would not generalize; a simpler check is to compare end-to-end latency with the re-noise term enabled versus disabled.
Extended reading notes
Core claim
At the core of Deltoris is the identity $y_t = W x_t = y_{t-1} + W (x_t - x_{t-1})$, which turns every linear projection in a diffusion-based VLA model into a differential update over the previous step's cached output. Because robot observations change little at 50–200 Hz, $\Delta x_t$ is sparse at the bit level, so a bit-serial accelerator that processes only the active 1-bits of the sign-magnitude representation of $\Delta x_t$ can skip over 90% of the arithmetic. The paper further claims that this bit-sparsity makes inference memory-bound rather than compute-bound, and that the proposed speculative inference—where a lightweight draft model proposes several future actions and the accurate large model verifies them in a batch—amortizes off-chip weight and activation traffic across control steps. Together with a 1D systolic bit-serial PE array that eliminates the workload imbalance of prior bit-serial designs, Deltoris claims up to 92.9% operation reduction, 34.2$\times$ speedup over a mobile GPU, 6.1$\times$ over the closest prior accelerator, and average success-rate loss of 0.2% on PAD, Diffusion Policy, and UVA.
Load-bearing premise
The speedup rests on the assumption that the numbers entering the network's layers change so little from one control step to the next that their bit-level difference is sparse; the paper demonstrates this for raw camera pixels, but not for internal activations after nonlinearities and after the stochastic re-noising term in each diffusion step.
Editorial extensions
If this is right
- If the claimed speedups hold, diffusion-based VLA policies like PAD, Diffusion Policy, and UVA could move from roughly 2 Hz on an edge SoC to the 50–200 Hz control rates required for stable closed-loop manipulation, without retraining or model distillation.
- The bit-sparsity algorithm is bit-accurate for linear operators, so the only accuracy loss comes from the speculative draft–verify step; this makes the approach usable in safety-sensitive control where approximate inference is not acceptable.
- Since temporal-aware bit-sparsity and inter-timestep similarity exploit orthogonal redundancies, a system combining both could push operation reduction beyond the 92.9% reported here for multi-step diffusion processes.
- The 1D systolic PE array design, if validated, removes the workload-imbalance bottleneck that has limited prior bit-serial accelerators, making bit-sparse execution viable at high utilization on other workloads with temporal redundancy.
Reading between the lines
- The paper's evidence for temporal similarity is at the pixel level; a direct layer-wise measurement of activation-difference sparsity on a real robot with the stochastic re-noise term active would show whether the 34.2$\times$ figure transfers beyond the three evaluated models and datasets.
- The batched-verification idea could extend beyond a single robot: if several policies share weights, the same amortization of weight and activation loads could apply across robots or tasks, not just across frames within one control loop.
- The graceful degradation seen at 2$\times$ scene acceleration suggests an adaptive policy: dynamically shrink the speculative window or fall back to full computation when motion estimates indicate that temporal similarity is dropping.
- The draft model uses roughly 1/5 of the denoising steps of the large model; using a single-step distilled draft could enlarge the speculative window and further amortize data loading, at the cost of a lower acceptance rate.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes Deltoris, an algorithm-hardware co-design framework for diffusion-based vision-language-action (VLA) inference. The algorithmic core observes that consecutive control steps in robotic loops are highly similar and rewrites each linear projection as y_t = y_{t-1} + W (x_t - x_{t-1}), so that only the bit-level differences in activations need to be computed. To mitigate the extra off-chip traffic caused by the differential computation, the paper introduces speculative inference, in which a small draft model generates several future control steps that are verified in batch by the full model. Finally, the paper presents a 1D systolic bit-serial PE array with weight sharing and claims up to 34.8x speedup over a mobile GPU and 6.1x over prior accelerators with only 0.2% average success-rate loss on PAD, DP, and UVA.
Significance. If the central temporal-sparsity premise holds, Deltoris would be a strong contribution: the differential update in Eq. (2) is an exact identity rather than an approximate fit, and the hardware evaluation is concrete, using post-layout RTL measurements, power simulation, and a cycle-accurate simulator. The approach also advances the architecture literature by targeting diffusion-based VLA inference rather than LLM or image-diffusion workloads. However, the load-bearing assumption that the tensors entering the linear layers are bit-sparse in their temporal difference is only demonstrated on raw camera pixels, not on the latent activations that Eq. (2) actually consumes. This gap, together with the need to clarify the verification protocol of speculative inference, is the main obstacle to accepting the paper's quantitative claims.
major comments (4)
- [Sec. 3, Sec. 4.1, Eq. (2), Fig. 8] The central operation-reduction claim rests on the assumption that Delta x_t entering each linear layer is bit-sparse. The only direct evidence of temporal similarity in Sec. 3 (Fig. 7) is measured on raw PushT pixels, not on the activations consumed by the linear projections in Eq. (2). Fig. 8 reports "bit sparsity" for PAD, DP, and UVA, but the text never states how it was measured, at which layers or tensors, or whether the stochastic sampler in Eq. (1) was used. This is load-bearing: Fig. 17c's 93% operation reduction and the speedup numbers scale with the sparsity of Delta x_t. Moreover, Eq. (1) contains the stochastic re-noising term sigma_t n'_t; if the initial latent or the re-noising noise is drawn independently per control step or per denoising step, the corresponding components of Delta x_t are differences of independent Gaussian draws, which are dense under the sign-magnitude representation used in Sec. 4.1. Please add per-layer activation-difference bit sparsity measurements for the actual inference schedule, including the first denoising step, and report the fraction of linear-layer input elements affected by freshly sampled noise.
- [Sec. 5, Fig. 12] The speculative-inference protocol as described verifies future actions using the draft model's predicted observations and states as inputs to the large model. In a closed-loop robot, the action at step t+1 should be a function of the actual observation received after executing the accepted action, not of the simulated observation generated by the draft model. The paper should state explicitly whether the success-rate evaluation in Fig. 16 uses this simulated-observation protocol, and it should provide a comparison against a variant in which the large model is re-run on the true observation before each action is executed. This is needed to distinguish a true verification scheme from a model-based rollout and to support the claim that only 0.2% success rate is lost.
- [Sec. 8.1, Fig. 16] The accuracy results are reported without confidence intervals or the number of evaluation episodes. The claim that Deltoris loses only 0.2% accuracy while the small model alone loses more than 7.7% is central to the paper's robustness argument, but with no variance information it is unclear whether 0.2% is within noise, especially given the small number of tasks and datasets. Please report per-task success rates, number of rollouts, standard errors, and the per-task acceptance rate of speculative candidates.
- [Sec. 7, Tbl. 2] The fairness of the hardware comparison needs clarification. The text states that the compute throughput of all accelerators is configured to be equivalent to 32x32 8-bit MAC units, but Deltoris is described as 128 PE arrays of 64 PEs at 1 GHz, while the baselines have very different organizations (e.g., 512 inner-product units for Pragmatic and 32x32 BitVert units for BBS). Please specify the exact per-cycle bit-level throughput, MAC count, and buffer partition for each design, and explain how the 2.45 mm2 area of Deltoris was obtained under the same technology assumptions. Otherwise the 6.1x speedup over prior accelerators may partly reflect configuration choices rather than the proposed techniques.
minor comments (5)
- [Abstract and Sec. 8.2] The abstract reports "up to 34.2x speedup over mobile GPUs," while Sec. 8.2 lists 34.8x, 46.1x, 24.4x, and 31.5x; the abstract also reports 822.0x energy savings while Fig. 17b lists 850x, 1040x, 746x, and 652x. These numbers should be reconciled.
- [Eq. (3)] In Eq. (3), v_{t,b} is called a "zero indicator," but the equation sums only over bits with nonzero v_{t,b}; please rename it an "active-bit indicator" to avoid confusion.
- [Fig. 8] Fig. 8 would benefit from a labeled y-axis, legend, and a precise definition of how bit sparsity is aggregated across layers, since the caption alone does not specify whether the numbers are weighted by tensor size or by FLOPs.
- [Sec. 7, speculative inference] The statement that the draft model "uses roughly 1/5 of the denoising computation of the original model" should specify the exact number of denoising steps and whether the draft and large models share weights or differ only in step count; this affects the interpretation of the speculative overhead.
- [Fig. 16] The rendered text in Fig. 16a appears garbled with numeric character references; please ensure the final figure is readable and the per-task success rates are legible.
Circularity Check
No significant circularity: temporal differencing is an exact algebraic identity and the headline speedups are measured in RTL/simulator evaluation, not fitted to match a target.
full rationale
Deltoris's central derivation is Eq. (2), y_t = y_{t-1} + W Δx_t, which is an exact algebraic identity; the differential product is then re-expressed in sign-magnitude bit-serial form in Eq. (3). Neither equation defines its claimed result in terms of itself, and no fitted parameter is relabeled as a prediction. The reported 92.9% operation reduction and 34.2x/6.1x speedups are presented as measured outcomes from RTL synthesis, post-layout simulation, and a cycle-accurate simulator on external benchmarks (Meta-World, PushT, LIBERO), not as consequences of constants chosen to force the headline numbers. Self-citations in the reference list (e.g., Astraea, Lumina, Cicero) are used for related-work context or general technique pointers and are not load-bearing; no uniqueness theorem or ansatz is imported from the authors' prior work to forbid alternatives. The main weaknesses of the paper—that activation-level bit sparsity behind the operation reduction is not fully specified, that the stochastic re-noising term in Eq. (1) could reduce temporal sparsity, and that the speculation threshold/window are tuned on PAD—are empirical validity and generalization risks, not circular reductions. Under the requirement that circularity be demonstrated by a quote and a specific reduction, no such step is present.
Assumptions & free parameters
free parameters (4)
- speculative acceptance threshold theta_th =
0.01
- speculative window size k =
3
- draft model denoising-step ratio =
1/5 of original denoising steps
- weight-sharing group size M =
16
assumptions (6)
- standard math Linearity of linear projections (W x_t = W x_{t-1} + W Delta_x_t) and exact bit-level accumulation of active bits.
- domain assumption Consecutive control-step observations and robot states are highly similar in real VLA deployments.
- domain assumption Latent activations in the denoising network inherit temporal similarity and remain bit-sparse after nonlinearities.
- domain assumption The reverse diffusion input at each control step can be treated as temporally similar.
- domain assumption A step-reduced draft model (1/5 denoising steps) produces candidate futures that the full model accepts often enough to amortize traffic.
- domain assumption Diffusion-based VLA inference is compute-bound before applying bit sparsity.
Cite this review
Pith. "Pith review of Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference." pith.science (2026). https://pith.science/paper/64OFD2TH
@misc{pith2026260804428,
author = {Pith},
title = {Pith review of: Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/64OFD2TH}},
note = {Machine review of arXiv:2608.04428}
}
abstract
Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. However, diffusion-based VLA models are compute-intensive and must run at high control frequency, e.g., 50-200 Hz. Thus, it imposes strict latency and energy constraints on edge devices. In this work, we present Deltoris, an algorithm-hardware co-design framework for efficient diffusion-based VLA inference. First, we exploit the temporal similarity of consecutive inputs and propose a \textit{temporal-aware bit-sparsity} algorithm that computes only the differences between consecutive inputs, eliminating redundant bit-level operations. To further address the extra off-chip traffic introduced by our algorithm, we propose a \textit{speculative inference} technique, which amortizes data loading across multiple control steps. Lastly, to support these techniques, we co-design a dedicated accelerator with customized 1D systolic bit-serial PE arrays that eliminate PE workload imbalance. Our evaluation shows that Deltoris achieves up to 34.2$\times$ speedup over mobile GPUs and 6.1$\times$ over prior accelerators, while maintaining comparable accuracy.
Figures
Figures from the paper (12 more)
Reference graph
Works this paper leans on
-
[1]
MLQ Agent. 2026. Boston Dynamics and Google DeepMind Form New AI Partnership for Humanoid Robots. https://mlq.ai/news/boston-dynamics- google-deepmind-form-new-ai-partnership-for-humanoid-robots/
2026
-
[2]
Jorge Albericio, Alberto Delmás, Patrick Judd, Sayeh Sharify, Gerard O’Leary, Roman Genov, and Andreas Moshovos. 2017. Bit-Pragmatic Deep Neural Net- work Computing. In2017 50th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 382–394
2017
-
[3]
Mohammad Bakhshalipour, Seyed Borna Ehsani, Mohamad Qadri, Dominic Guri, Maxim Likhachev, and Phillip B. Gibbons. 2022. RACOD: algorithm/hardware co-design for mobile robot path planning. InProceedings of the 49th Annual International Symposium on Computer Architecture(New York, New York)(ISCA ’22). Association for Computing Machinery, New York, NY, USA, ...
arXiv 2022
-
[5]
Suneel Belkhale, Tianli Ding, Ted Xiao, Pierre Sermanet, Quon Vuong, Jonathan Tompson, Yevgen Chebotar, Debidatta Dwibedi, and Dorsa Sadigh. 2024. RT-H: Action Hierarchies using Language. (2024)
2024
-
[6]
2026.𝜋 0: A Vision- Language-Action Flow Model for General Robot Control.arXiv(2026)
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, James Tanner, Quan Vuong, Anna Walling, Haohuan Wang, and Ury Zhilinsky. 2026.𝜋 0: ...
2026
-
[7]
Mark Buckler, Philip Bedoukian, Suren Jayasuriya, and Adrian Sampson
-
[8]
Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, and Kai-Chiang Wu
-
[9]
Peiqing Chen, Minghao Li, Zishen Wan, Yu-Shun Hsiao, Minlan Yu, Vijay Janapa Reddi, and Zaoxing Liu. 2025. OctoCache: Caching Voxels for Accelerating 3D Occupancy Mapping in Autonomous Systems. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. Association for Computin...
arXiv 2025
Show all 117 references
-
[10]
Pengtao Chen, Mingzhu Shen, Peng Ye, Jianjian Cao, Chongjun Tu, Christos- Savvas Bouganis, Yiren Zhao, and Tao Chen. 2024. Delta-DiT: A Training- Free Acceleration Method Tailored for Diffusion Transformers.arXiv preprint arXiv:2406.01125(2024)
2024 arXiv
-
[12]
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. 2023. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. InProceedings of Robotics: Science and Systems (RSS)
2023
-
[13]
Wonkyo Choe, Rongxiang Wang, and Felix Xiaozhu Lin. 2025. AnA: An Attentive Autonomous Driving System. InProceedings of the 30th ACM In- ternational Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1(Rotterdam, Netherlands)(ASPLOS ’25...
2025
-
[14]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah
-
[15]
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022. Flashat- tention: Fast and memory-efficient exact attention with io-awareness.Advances in Neural Information Processing Systems35 (2022), 16344–16359
2022
-
[17]
Alberto Delmas Lascorz, Patrick Judd, Dylan Malone Stuart, Zissis Poulos, Mostafa Mahmoud, Sayeh Sharify, Milos Nikolic, Kevin Siu, and Andreas Moshovos. 2019. Bit-Tactical: A Software/Hardware Approach to Exploiting Value and Bit Sparsity in Neural Networks. InProceedings of ...
2019
-
[18]
Li, Xue Lin, Zhenman Fang, and Yanzhi Wang
Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Li-Yu Daisy Liu, Zhenglun Kong, Xin Meng, Z. Li, Xue Lin, Zhenman Fang, and Yanzhi Wang. 2023. HeatViT: Hardware-Efficient Adaptive Token Pruning for Vision Transformers.2023 IEEE 13 Conference’17, July 2017, Washington, DC, USA Z...
2023
-
[19]
Shichen Dong, Wen Cheng, Jiayu Qin, and Wei Wang. 2024. QAQ: Quality Adaptive Quantization for LLM KV Cache. (2024). arXiv:2403.04643 [cs.CL]
2024 arXiv
-
[20]
Yilun Du, Mengjiao Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B Tenen- baum, Dale Schuurmans, and Pieter Abbeel. 2023. Learning Universal Policies via Text-Guided Video Generation. In2023 Conference on Neural Information Processing Systems
2023
-
[21]
Yu Feng, Patrick Hansen, Paul N Whatmough, Guoyu Lu, and Yuhao Zhu. 2023. Fast and Accurate: Video Enhancement Using Sparse Depth. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 4492–4500
2023
-
[22]
Yu Feng, Weikai Lin, Yuge Cheng, Zihan Liu, Jingwen Leng, Minyi Guo, Chen Chen, Shixuan Sun, and Yuhao Zhu. 2025. Lumina: Real-Time Mobile Neural Rendering by Exploiting Computational Redundancy. InProceedings of the 52nd Annual International Symposium on Computer Architecture
2025
-
[24]
Yu Feng, Paul Whatmough, and Yuhao Zhu. 2019. Asv: Accelerated stereo vision system. InProceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture. 643–656
2019
-
[25]
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers. arXiv preprint arXiv:2210.17323(2022)
2022 arXiv
-
[26]
Yanjiang Guo, Yucheng Hu, Jianke Zhang, Yen-Jen Wang, Xiaoyu Chen, Chaochao Lu, and Jianyu Chen. 2024. Prediction with Action: Visual Policy Learning via Joint Denoising Process. In2024 Conference on Neural Information Processing Systems
2024
-
[27]
gym pusht. 2023. A gymnasium environment PushT dataset. https://github. com/huggingface/gym-pusht/
2023
-
[28]
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, Poriya Panet, Sapir Weissbuch, Victor Kulikov, Yaki Bitterman, Zeev Melumian, and Ofir Bibi. 2024. LTX-Video: Realtime Video Latent Diffus...
2024 arXiv
-
[31]
Jaehoon Heo, Adiwena Putra, Jieon Yoon, Sungwoong Yune, Hangyeol Lee, Ji-Hoon Kim, and Joo-Young Kim. 2025. EXION: Exploiting Inter-and Intra- Iteration Output Sparsity for Diffusion Models. In2025 IEEE International Sym- posium on High Performance Computer Architecture (HPCA)...
2025
-
[32]
Alec Hively. 2026. Tesla Optimus Robot’s Public Failure Has To Be Seen To Be Believed. https://www.bgr.com/2065073/tesla-optimus-robot-public-failure/
2026
-
[33]
Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion proba- bilistic models.Advances in neural information processing systems33 (2020), 6840–6851
2020
-
[34]
Susan Hong. 2026. Humanoid Robots Exit Labs: Mapping the Technical Path to Embodied AI at AW 2026. https://www.eetimes.com/humanoid-robots-exit- labs-mapping-the-technical-path-to-embodied-ai-at-aw-2026/
2026
-
[35]
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024. KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization.arXiv preprint arXiv:2401.18079(2024)
2024 arXiv
-
[36]
Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen. 2024. Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations. In Forty-Second International Conference on Machin...
2024
-
[37]
Xiaotong Huang, He Zhu, Tianrui Ma, Yuxiang Xiong, Fangxin Liu, Zhezhi He, Yiming Gan, Zihan Liu, Jingwen Leng, Yu Feng, and Minyi Guo. 2026. Splatonic: Architectural Support for 3D Gaussian Splatting SLAM via Sparse Processing. InProceedings of the 32nd IEEE International Sym...
2026
-
[39]
Mingxiao Huo, Jiayi Zhang, Hewei Wang, Jinfeng Xu, Zheyu Chen, Huilin Tai, and Yijun Chen. 2025. Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding.arXiv preprint arXiv:2509.11961 (2025)
2025
-
[40]
Kha Minh Huynh, Thien Thanh Bui Nguyen, Hai Vu Nguyen, Khoa Dac Tran, Kenichi Iwata, Katsuya Mizumoto, Nobuhiko Honda, Keisuke Matsumoto, Kat- sushige Matsubara, and Seiji Mochizuki. 2017. 16.8 GB/s LPDDR4-3200@32-bit memory access bandwidth. In2017 7th International Conferenc...
2017
-
[41]
Micron Technology Inc. 2025. Micron System Power Calculators. https://www. micron.com/support/tools-and-utilities/power-calc
2025
-
[42]
Aamodt, and Andreas Moshovos
Patrick Judd, Jorge Albericio, Tayler Hetherington, Tor M. Aamodt, and Andreas Moshovos. 2016. Stripes: Bit-serial deep neural network computing. In2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). 1–12. https://doi.org/10.1109/MICRO.2016.7783722
2016
-
[43]
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. 2024. O...
2024 arXiv
-
[44]
Moo Jin Kim, Chelsea Finn, and Percy Liang. 2025. Fine-Tuning Vision-Language- Action Models: Optimizing Speed and Success.arXiv preprint arXiv:2502.19645 (2025)
2025 arXiv
-
[45]
Sungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park, and Won Woo Ro
-
[46]
Weihao Kong, Yifan Hao, Qi Guo, Yongwei Zhao, Xinkai Song, Xiaqing Li, Mo Zou, Zidong Du, Rui Zhang, Chang Liu, et al. 2024. Cambricon-d: Full-network differential acceleration for diffusion models. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (...
2024
-
[47]
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. 2024. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603 (2024)
2024 arXiv
-
[48]
Srivatsan Krishnan, Zishen Wan, Kshitij Bhardwaj, Paul Whatmough, Aleksan- dra Faust, Sabrina Neuman, Gu-Yeon Wei, David Brooks, and Vijay Janapa Reddi
-
[49]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP ’23). Association for Computing Machinery, New York, NY, USA...
2023
-
[50]
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. InProceedings of the 40th International Conference on Machine Learning(Honolulu, Hawaii, USA)(ICML’23). JMLR.org, Article 795, 13 pages
2023
-
[51]
Jianze Li, Jiezhang Cao, Zichen Zou, Xiongfei Su, Xin Yuan, Yulun Zhang, Yong Guo, and Xiaokang Yang. 2025. Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator. (2025). arXiv:2410.04224 [cs.CV] https://arxiv.org/abs...
2025 arXiv
-
[52]
Lijiang Li, Huixia Li, Xiawu Zheng, Jie Wu, Xuefeng Xiao, Rui Wang, Min Zheng, Xin Pan, Fei Chao, and Rongrong Ji. 2023. Autodiffusion: Training- free optimization of time steps and architectures for automated diffusion model acceleration. InProceedings of the IEEE/CVF Interna...
2023
-
[53]
InProceedings of the 55th Annual IEEE/ACM International Sym- posium on Microarchitecture(Chicago, Illinois, USA)(MICRO ’22)
Automatic Domain-Specific SoC Design for Autonomous Unmanned Aerial Vehicles. InProceedings of the 55th Annual IEEE/ACM International Sym- posium on Microarchitecture(Chicago, Illinois, USA)(MICRO ’22). IEEE Press, 300–317. https://doi.org/10.1109/MICRO56248.2022.00033
-
[54]
Xinghang Li, Minghuan Liu, Hanbo Zhang, Cunjun Yu, Jie Xu, Hongtao Wu, Chilam Cheang, Ya Jing, Weinan Zhang, Huaping Liu, Hang Li, and Tao Kong
-
[55]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration. InMLSys
2024
-
[56]
Yoseph Linde, Andres Buzo, and Robert Gray. 1980. An algorithm for vector quantizer design.IEEE Transactions on Communications28, 1 (1980), 84–95
1980
-
[57]
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. 2023. LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot 14 Deltoris: Real-time VLA Inference in Embodied AI Conference’17, July 2017, Washington, DC, USA Learning.arXiv preprint arXiv:...
2023 arXiv
-
[58]
Shuang Li, Yihuai Gao, Dorsa Sadigh, and Shuran Song. 2025. Unified Video Action Model. InProceedings of Robotics: Science and Systems (RSS)
2025
-
[59]
Joseph Liu, Joshua Geddes, Ziyu Guo, Haomiao Jiang, and Mahesh Kumar Nandwana. 2024. SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers.arXiv preprint arXiv:2411.10510(2024)
2024 arXiv
-
[60]
Vision-Language Foundation Models as Effective Robot Imitators.arXiv preprint arXiv:2311.01378(2023)
2023 arXiv
-
[61]
Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. 2024. RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.arXiv preprint arXiv:2410.07864(2024)
2024 arXiv
-
[62]
Weizhuang Liu, Bo Yu, Yiming Gan, Qiang Liu, Jie Tang, Shaoshan Liu, and Yuhao Zhu. 2021. Archytas: A Framework for Synthesizing and Dynami- cally Optimizing Accelerators for Robotic Localization. InMICRO-54: 54th An- nual IEEE/ACM International Symposium on Microarchitecture(...
2021
-
[63]
Xiaoxuan Liu, Lanxiang Hu, Peter Bailis, Alvin Cheung, Zhijie Deng, Ion Sto- ica, and Hao Zhang. 2024. Online speculative decoding. InProceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 1258, 16 pages
2024
-
[64]
Haosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu, Aiyue Chen, Jing Lin, Yiwu Yao, Chen Chen, Jingwen Leng, Yu Feng, and Minyi Guo. 2025. Astraea: A Token-wise Acceleration Framework for Video Diffusion Transformers. (2025). arXiv:2506.05096 [cs.CV] https://arxiv.org/abs/2506.05096
2025
-
[65]
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2024. Deepcache: Accelerating diffusion models for free. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15762–15772
2024
-
[66]
Peiqi Liu, Yaswanth Orru, Chris Paxton, Nur Muhammad Mahi Shafiullah, and Lerrel Pinto. 2024. OK-Robot: What Really Matters in Integrating Open- Knowledge Models for Robotics.arXiv preprint arXiv:2401.12202(2024)
2024 arXiv
-
[67]
Neuman, Radhika Ghosal, Thomas Bourgeat, Brian Plancher, and Vijay Janapa Reddi
Sabrina M. Neuman, Radhika Ghosal, Thomas Bourgeat, Brian Plancher, and Vijay Janapa Reddi. 2023. RoboShape: Using Topology Patterns to Scalably and Flexibly Deploy Accelerators Across Robots. InProceedings of the 50th Annual International Symposium on Computer Architecture(Or...
2023
-
[68]
Neuman, Brian Plancher, Thomas Bourgeat, Thierry Tambe, Srinivas Devadas, and Vijay Janapa Reddi
Sabrina M. Neuman, Brian Plancher, Thomas Bourgeat, Thierry Tambe, Srinivas Devadas, and Vijay Janapa Reddi. 2021. Robomorphic computing: a design methodology for domain-specific accelerators parameterized by robot morphol- ogy. InProceedings of the 26th ACM International Conf...
2021
-
[69]
Dima Nikiforov, Shengjun Chris Dong, Chengyi Lux Zhang, Seah Kim, Borivoje Nikolic, and Yakun Sophia Shao. 2023. RoSÉ: A Hardware-Software Co- Simulation Infrastructure Enabling Pre-Silicon Full-Stack Robotics SoC Evalua- tion. InProceedings of the 50th Annual International Sy...
2023 doi
-
[70]
Hang Lu, Liang Chang, Chenglong Li, Zixuan Zhu, Shengjian Lu, Yanhuan Liu, and Mingzhe Zhang. 2021. Distilling Bit-level Sparsity Parallelism for General Purpose Deep Learning Acceleration. InMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO ’2...
2021 doi
-
[71]
Nvidia. 2025. NVIDIA Jetson Thor: The ultimate platform for physical AI and robotics. https://www.nvidia.com/en-us/autonomous-machines/embedded- systems/jetson-thor/
2025
-
[72]
Zentner, Ryan Julian, J K Terry, Isaac Woun- gang, Nariman Farsad, and Pablo Samuel Castro
Reginald McLean, Evangelos Chatzaroulas, Luc McCutcheon, Frank Röder, Tianhe Yu, Zhanpeng He, K.R. Zentner, Ryan Julian, J K Terry, Isaac Woun- gang, Nariman Farsad, and Pablo Samuel Castro. 2025. Meta-World+: An Improved, Standardized, RL Benchmark. InThe Thirty-ninth Annual ...
2025
-
[73]
Tim Salimans and Jonathan Ho. 2022. Progressive Distillation for Fast Sam- pling of Diffusion Models. Inthe Tenth International Conference on Learning Representations (2022)
2022
-
[74]
Satyabrata Sarangi and Bevan Baas. 2021. DeepScaleTool: A tool for the accu- rate estimation of technology scaling in the deep-submicron era. In2021 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 1–5
2021
-
[75]
Deval Shah and Tor M. Aamodt. 2024. Collision Prediction for Robotics Ac- celerators. InProceedings of the 51st Annual International Symposium on Com- puter Architecture(Buenos Aires, Argentina)(ISCA ’24). IEEE Press, 566–581. https://doi.org/10.1109/ISCA59077.2024.00048
2024
-
[76]
Nvidia. 2023. Jetson Orin for Next-Gen Robotics. https://www.nvidia.com/en- us/autonomous-machines/embedded-systems/jetson-orin/
2023
-
[77]
Sayeh Sharify, Alberto Delmas Lascorz, Mostafa Mahmoud, Milos Nikolic, Kevin Siu, Dylan Malone Stuart, Zissis Poulos, and Andreas Moshovos. 2019. La- conic Deep Learning Inference Acceleration. In2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)...
2019
-
[78]
Lawson, Behnam Khaleghi, and Hadi Esmaeilzadeh
Jacob Sacks, Divya Mahajan, Richard C. Lawson, Behnam Khaleghi, and Hadi Esmaeilzadeh. 2018. Robox: an end-to-end solution to accelerate autonomous control in robotics. InProceedings of the 45th Annual International Symposium on Computer Architecture(Los Angeles, California)(I...
2018
-
[79]
Man Shi, Vikram Jain, Antony Joseph, Maurice Meijer, and Marian Verhelst. 2024. BitWave: Exploiting Column-Based Bit-Level Sparsity for Deep Learning Accel- eration. In2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 732–746. https://doi.org/...
2024
-
[80]
Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, and Xiaoyao Liang. 2024. CMC: Video Transformer Acceleration via CODEC Assisted Matrix Condensing. InProceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating System...
2024
-
[81]
Zhuoran Song, Feiyang Wu, Xueyuan Liu, Jing Ke, Naifeng Jing, and Xiaoyao Liang. 2020. Vr-dann: Real-time video recognition via decoder-assisted neural network acceleration. In2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 698–710
2020
-
[82]
Deval Shah, Ningfeng Yang, and Tor M. Aamodt. 2023. Energy-Efficient Realtime Motion Planning. InProceedings of the 50th Annual International Symposium on Computer Architecture(Orlando, FL, USA)(ISCA ’23). Association for Computing Machinery, New York, NY, USA, Article 57, 17 ...
2023
-
[83]
Justin Ting, Minsik Kim, Junkang Zhu, Haotian Sheng, and Zhengya Zhang
-
[84]
Hardik Sharma, Jongse Park, Naveen Suda, Liangzhen Lai, Benson Chau, Joon Kyung Kim, Vikas Chandra, and Hadi Esmaeilzadeh. 2018. Bit Fusion: Bit-Level Dynamically Composable Architecture for Accelerating Deep Neural Network. In2018 ACM/IEEE 45th Annual International Symposium ...
2018
-
[85]
Zishen Wan, Yuhang Du, Mohamed Ibrahim, Jiayi Qian, Jason Jabbour, Yang (Katie) Zhao, Tushar Krishna, Arijit Raychowdhury, and Vijay Janapa Reddi. 2025. ReCA: Integrated Acceleration for Real-Time and Efficient Coopera- tive Embodied Autonomous Agents. InProceedings of the 30t...
2025
-
[86]
Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jin- gren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pandeng Li, P...
-
[87]
Songsheng Wang, Rucheng Yu, Zhihang Yuan, Chao Yu, Feng Gao, Yu Wang, and Derek F Wong. 2025. Spec-vla: speculative decoding for vision-language- action models with relaxed acceptance. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 26...
2025
-
[88]
Aaron Stillmaker and Bevan Baas. 2017. Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm.Integration58 (2017), 74–81
2017
-
[91]
Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess
Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springenberg, Sergey Levine, Chelsea Finn, and Danny Driess. 2026. MEM: Multi-Scale ...
2026
-
[92]
Junjie Wen, Yichen Zhu, Zhibing Tang, Jinming Li, Yaxin Peng, Chaomin Shen, and Feifei Feng. 2025. DexVLA: Vision-Language Model with Plug-In Diffusion Expert for Visuomotor Policy Learning.arXiv preprint arXiv:2502.05855(2025)
2025 arXiv
-
[93]
Felix Wimbauer, Bichen Wu, Edgar Schoenfeld, Xiaoliang Dai, Ji Hou, Zijian He, Artsiom Sanakoyeu, Peizhao Zhang, Sam Tsai, Jonas Kohler, et al. 2024. Cache me if you can: Accelerating diffusion models through block caching. InProceedings of the IEEE/CVF Conference on Computer ...
2024
-
[94]
Wan: Open and Advanced Large-Scale Video Generative Models.arXiv preprint arXiv:2503.20314(2025)
2025 arXiv
-
[95]
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. InProceedings of the 40th International Conference on Machine Learning
2023
-
[96]
Yuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang, Yingyan Li, Yuntao Chen, Xinlong Wang, and Zhaoxiang Zhang. 2025. Unified Vision-Language- Action Model.arXiv preprint arXiv:2506.19850(2025)
2025 arXiv
-
[97]
Siyu Xu, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang, and Chang Xu
-
[98]
Junjie Wen, Minjie Zhu, Yichen Zhu, Zhibin Tang, Jinming Li, Zhongyi Zhou, Chengmeng Li, Xiaoyu Liu, Yaxin Peng, Chaomin Shen, and Feifei Feng. 2024. DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression.Forty-Second International Conference on...
2024
-
[99]
Junjie Wen, Yichen Zhu, Jinming Li, Minjie Zhu, Kun Wu, Zhiyuan Xu, Ning Liu, Ran Cheng, Chaomin Shen, Yaxin Peng, Feifei Feng, and Jian Tang. 2025. TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation. (2025). arXiv:2409.12514 [cs.RO] h...
2025 arXiv
-
[100]
Yuxin Yang, Xiaoming Chen, and Yinhe Han. 2023. Dadu-RBD: Robot Rigid Body Dynamics Accelerator with Multifunctional Pipelines. InProceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture(Toronto, ON, Canada)(MICRO ’23). Association for Computing Mac...
2023
-
[101]
Ziyu Ying, Shulin Zhao, Haibo Zhang, Cyan Subhra Mishra, Sandeepa Bhuyan, Mahmut T Kandemir, Anand Sivasubramaniam, and Chita R Das. 2022. Ex- ploiting Frame Similarity for Efficient Inference on Edge Devices. In2022 IEEE 42nd International Conference on Distributed Computing ...
2022
-
[102]
Haocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu, Muyang Li, Xiuyu Li, Yujun Lin, Han Cai, Jintao Zhang, Dacheng Li, et al. 2025. Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity. arXiv preprint arXiv:2502.01776(2025)
2025 arXiv
-
[103]
Haoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu, Yang Zhao, Yongan Zhang, Chaojian Li, Baopu Li, and Yingyan Lin. 2023. ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-Design. In2023 IEEE International Symposium on High-Performance Computer ...
2023
-
[104]
Mengda Xu, Zhenjia Xu, Yinghao Xu, Cheng Chi, Gordon Wetzstein, Manuela Veloso, and Shuran Song. 2024. Flow as the Cross-domain Manipulation Interface. In8th Annual Conference on Robot Learning. https://openreview.net/forum?id= cNI0ZkK1yC
2024
-
[105]
Jianke Zhang, Yanjiang Guo, Yucheng Hu, Xiaoyu Chen, Xiang Zhu, and Jianyu Chen. 2025. UP-VLA: A Unified Understanding and Prediction Model for Em- bodied Agent.arXiv preprint arXiv:2501.18867(2025)
2025 arXiv
-
[106]
Vla-cache: Towards efficient vision-language-action model via adaptive token caching in robotic manipulation.arXiv e-prints(2025), arXiv–2502
2025
-
[107]
Kashu Yamazaki, Viet-Khoa Vo-Ho, Darshan Bulsara, and Ngan Le. 2022. Spiking neural networks and their applications: A review.Brain sciences12, 7 (2022), 863
2022
-
[108]
Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion models: A comprehensive survey of methods and applications.Comput. Surveys56, 4 (2023), 1–39
2023
-
[109]
Shulin Zhao, Haibo Zhang, Cyan Subhra Mishra, Sandeepa Bhuyan, Ziyu Ying, Mahmut Taylan Kandemir, Anand Sivasubramaniam, and Chita Das. 2021. Holoar: On-the-fly optimization of 3d holographic processing for augmented reality. InMICRO-54: 54th Annual IEEE/ACM International Symp...
2021
-
[110]
Xuanlei Zhao, Xiaolong Jin, Kai Wang, and Yang You. 2024. Real-time video generation with pyramid attention broadcast.arXiv preprint arXiv:2408.12588 (2024)
2024 arXiv
-
[111]
Seungjae Yoo, Hangyeol Kim, and Joo-Young Kim. 2024. AdapTiV: Sign- Similarity Based Image-Adaptive Token Merging for Vision Transformer Accel- eration. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). 64–77. https://doi.org/10.1109/MICRO61859.2024.00015
2024
-
[112]
Zihao Zheng, Zhihao Mao, Sicheng Tian, Maoliang Li, Jiayu Chen, Xinhao Sun, Zhaobo Zhang, Xuanzhe Liu, Donggang Cao, Hong Mei, et al . 2026. Heisd: Hybrid speculative decoding for embodied vision-language-action models with kinematic awareness.arXiv preprint arXiv:2603.17573(2026)
2026 arXiv
-
[113]
Zhang Yushuo, Liu Jia, and Qiao Xinyi. 2026. China’s Embodied-AI Firms Exhibit Dancing, Fighting Robots at CES. https://www.yicaiglobal.com/news/in- photos-chinese-robots-shine-at-ces
2026
-
[114]
He Zhu, Zheng Liu, Xingyang Li, Anbang Wu, Jieru Zhao, Fangxin Liu, Yiming Gan, Jingwen Leng, and Yu Feng. 2026. Nebula: Infinite-Scale 3D Gaussian Splatting in VR via Collaborative Rendering and Accelerated Stereo Rasteriza- tion(ASPLOS ’26). Association for Computing Machine...
2026
-
[115]
Jintao Zhang, Kaiwen Zheng, Kai Jiang, Haoxu Wang, Ion Stoica, Joseph E Gonzalez, Jianfei Chen, and Jun Zhu. 2025. TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times.arXiv preprint arXiv:2512.16093(2025)
2025
-
[116]
Suquan Zhang, Yu Hu, Yunfei Xiang, Dawei Zhao, Yuanfan Xu, Qingmin Liao, Jincheng Yu, and Yu Wang. 2025. IDEA-GP: Instruction-Driven Architecture with Efficient Online Workload Allocation for Geometric Perception. InProceedings of the 52nd Annual International Symposium on Com...
2025
-
[117]
Shulin Zhao, Haibo Zhang, Sandeepa Bhuyan, Cyan Subhra Mishra, Ziyu Ying, Mahmut T Kandemir, Anand Sivasubramaniam, and Chita R Das. 2020. Déja view: Spatio-temporal compute reuse for ‘energy-efficient 360 vr video stream- ing. In2020 ACM/IEEE 47th Annual International Symposi...
2020
-
[120]
Zihao Zheng, Zhihao Mao, Maoliang Li, Jiayu Chen, Xinhao Sun, Zhaobo Zhang, Donggang Cao, Hong Mei, and Xiang Chen. 2026. Kerv: Kinematic-rectified speculative decoding for embodied vla models.arXiv preprint arXiv:2603.01581 (2026)
2026 arXiv
-
[122]
Peiyuan Zhi, Zhiyuan Zhang, Yu Zhao, Muzhi Han, Zeyu Zhang, Zhitian Li, Ziyuan Jiao, Baoxiong Jia, and Siyuan Huang. 2025. Closed-Loop Open- Vocabulary Mobile Manipulation with GPT-4V. In2025 International Conference on Robotics and Automation (ICRA)
2025
-
[124]
Yuhao Zhu, Anand Samajdar, Matthew Mattina, and Paul Whatmough. 2018. Euphrates: Algorithm-soc co-design for low-power mobile continuous vision. arXiv preprint arXiv:1803.11232(2018)
2018 arXiv
-
[125]
Chang Zou, Xuyang Liu, Ting Liu, Siteng Huang, and Linfeng Zhang. 2024. Accelerating diffusion transformers with token-wise feature caching.arXiv preprint arXiv:2410.05317(2024)
2024 arXiv
-
[126]
Pavlo Zvenyhorodskyi and Scott Singer. 2025. Embodied AI: China’s Big Bet on Smart Robots. https://carnegieendowment.org/research/2025/11/embodied-ai- china-smart-robots 16
2025
-
[2018]
In2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA)
EVA2: Exploiting temporal redundancy in live computer vision. In2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). IEEE, 533–546
-
[2023]
Diffusion models in vision: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 9 (2023), 10850–10869
2023
-
[2024]
Palu: Compressing KV-Cache with Low-Rank Projection. (2024). arXiv:2407.21118 [cs.AI] https://arxiv.org/abs/2407.21118
2024 arXiv
-
[2025]
In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA)
Ditto: Accelerating Diffusion Model via Temporal Value Similarity. In 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 338–352
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.