REVIEW 3 major objections 4 minor 82 references
A 0.9-billion-parameter robot model beats 7B baselines on all four LIBERO-Plus suites by structuring supervision instead of adding scale.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 02:15 UTC pith:L7FPV646
load-bearing objection Strong internal evidence and honest reporting, but the headline 0.9B-beats-7B claim is missing the one sentence that makes it fair — what the baselines trained on. the 3 major comments →
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper establishes that a sub-billion-parameter VLA can lead all published LIBERO-Plus baselines, including the strongest 7B systems, when supervision is deliberately structured along the axes the benchmark perturbs. The three components are not interchangeable: removing the paraphrase augmentation drops the Language Instructions axis by 12.3 points while leaving other axes unchanged; removing chain-of-thought distillation costs most on physical-state perturbations like Robot Initial States; and removing temporal history hurts most under kinematic perturbation. The reasoning structure is causally load-bearing: at test time, replacing the episode Plan with an empty or contradictory span co
What carries the argument
The central mechanism is hierarchical chain-of-thought distillation combined with dual-view temporal input. A 35B vision-language teacher generates an episode-level Plan and a chunk-level Think span with three structured slots: Phase, Gripper, and Next. The 0.9B student is trained to predict both these reasoning tokens and the action chunk from a stream of eight third-person and eight wrist frames with textual frame-index markers, plus an 8-dimensional proprioception vector and a paraphrased instruction. The Plan is generated once per episode and can be cached, halving steady-state inference cost; the Think span acts as an auxiliary structured prediction target that conditions the action hea
Load-bearing premise
The comparison is fair only if the published 7B baselines had access to the same perturbed LIBERO-Plus training data that CoTinyVLA was fine-tuned on; the paper never states what training data the baselines used.
What would settle it
Train the strongest 7B baseline (OpenVLA-OFT+) on the union of the four LIBERO-Plus suites for two epochs with the same inference budget and evaluation protocol; if the 4.7 to 15.9 point margins disappear or invert, the reported advantage comes from training on the evaluation distribution rather than from structured supervision at 0.9B. A second, complementary test: retrain CoTinyVLA on unperturbed LIBERO only and measure LIBERO-Plus success; a large drop would confirm the method depends on seeing perturbed tasks during training.
If this is right
- If the claim holds, large backbones are not necessary for robustness on this benchmark; a sub-billion model with structured supervision can be the leading policy, which directly addresses memory-constrained deployment.
- The perturbation-axis decomposition means each design component addresses a specific failure source: paraphrase augmentation for linguistic variation, temporal history for kinematic perturbation, and chain-of-thought for physical-state shifts.
- The episode Plan is not a decorative annotation; its content enters the action representation, so making the plan cacheable at inference time is a safe saving that reduces latency by about half without changing success.
- The measured 2.25GiB footprint and 1.37s per chunk on an L40S provide a concrete operating point for robots with embedded GPU budgets.
- On standard LIBERO the model averages 97.5%, matching the strongest 7B baseline and confirming that robustness gains do not come at the cost of in-distribution capability.
Where Pith is reading between the lines
- The most direct test of the paper's causal claim would be to train a 7B baseline on the same union of the four LIBERO-Plus suites with the same two-epoch recipe; the paper does not state what training data the published baselines used, so part of the margin could reflect training on the evaluation distribution rather than structured supervision substituting for scale.
- The teacher-dependency cost is not counted in the reported training budget; if a weaker teacher substantially degrades the Think labels, the method's practical gains may rely on a 35B teacher that many robotics teams cannot afford to run.
- The reasoning span is the dominant inference cost, so the approach composes with any decoding-side speedup (speculative decoding, early exit, or overlapping generation with execution), which the paper explicitly leaves untested on physical robots.
- The largest remaining gap is the Long suite under combined long-horizon execution and instruction variation; a natural extension would be to increase history length or add a memory mechanism specifically for that regime.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CoTinyVLA, a ~0.9B-parameter vision-language-action model built on a Qwen3.5-0.8B backbone, with three proposed components: dual-view temporal input (eight third-person and eight wrist frames with textual markers), hierarchical chain-of-thought distillation (episode-level Plan and chunk-level Think), and paraphrase augmentation for the 40 base instructions. It reports state-of-the-art results on the LIBERO-Plus robustness benchmark, exceeding the strongest 7B baselines on all four suites by 4.7, 2.8, 15.9, and 3.0 points, with reconstructed confidence intervals excluding zero, at a measured 2.25 GiB peak inference memory. The paper also provides axis-separated ablations, test-time interventions, and cost profiling to support the claim that structured supervision can substitute for model scale.
Significance. If the cross-model comparison is fair, this is a practically significant result: it suggests that a sub-billion-parameter VLA can exceed 3B–7B baselines on a robustness benchmark, with a detailed and credible resource characterization. The manuscript has clear strengths: it ships code, reconstructs confidence intervals transparently, uses paired test-time interventions, and separates ablation effects by perturbation axis. The central uncertainty is whether the reported margins reflect the proposed components or an asymmetry in training data access between CoTinyVLA and the published baselines.
major comments (3)
- [§3.5, §4.1–4.2, Tables 1–5] The paper fine-tunes CoTinyVLA on the union of the four LIBERO-Plus suites (§3.5) and evaluates on LIBERO-Plus (§4.1), but it never states what training data the published baselines used. If the baselines were trained on the original LIBERO tasks while CoTinyVLA was trained on the 10,030 LIBERO-Plus perturbed tasks (which are modifications of the same 40 base tasks), then CoTinyVLA has been trained on the evaluation distribution. This would give it a structural advantage unrelated to dual-view input, CoT distillation, or paraphrase augmentation, and would undermine the headline claim that structured supervision lets a 0.9B model beat 7B models. Please state the baseline training protocols explicitly, and report a CoTinyVLA variant trained only on original LIBERO (or otherwise matched to the baselines' training data). Without this, the cross-model margin cannot be attributed to the propos
- [§4.5, Table 3] The standard LIBERO comparison is affected by the same training-data asymmetry. CoTinyVLA is trained on LIBERO-Plus, which includes perturbed variants of the same 40 base task families used in standard LIBERO. Its standard LIBERO average of 97.5% therefore reflects exposure to augmented versions of the evaluation tasks, whereas the baselines were trained on the original LIBERO data. The claim that CoTinyVLA matches RIPT-VLA on standard LIBERO should be framed in light of this asymmetry, or supported by a variant trained on the original LIBERO benchmark.
- [§4.1] The paper refers to 'the standard LIBERO-Plus protocol' but does not define the benchmark's intended train/test split. A benchmark protocol is not fully specified by the evaluation rule that each task is rolled out once. Please clarify whether the LIBERO-Plus benchmark permits training on the perturbed test tasks, and how the published baseline numbers were produced in this respect. This is load-bearing for the main comparison, not a presentation detail.
minor comments (4)
- [Table 2] The text says CoTinyVLA leads on 'five of the seven perturbation axes' on Object, but according to Table 2 it leads on six axes (all except Language Instructions, where OpenVLA-OFT+ scores 100.0 versus 89.8).
- [Appendix G / Table 12] The ablation sweep is trained for one epoch whereas the headline model is trained for two epochs. The one-epoch reference is clearly labeled, but this fact should appear in the main text near the ablation discussion, not only in the supplementary training setup, to avoid over-reading the absolute levels in Table 12.
- [Table 16] The row label 'Most recent frame repeated16×' is ambiguous; clarify that it means all 16 history slots (eight per view) are filled with the most recent frame, not that the window is extended to 16 repetitions per view.
- [§3.1] The paper says the model has 'approximately 0.9B' parameters. Please report the exact parameter count or a breakdown by backbone, added embeddings, projection, and action head.
Circularity Check
No significant circularity: the central LIBERO-Plus result is an externally benchmarked empirical claim, with one minor self-referential validation probe.
specific steps
-
fitted input called prediction
[Sec. 3.3 (teacher gripper hint) and Sec. L / Table 17 (gripper agreement check)]
"The teacher prompt for Think reads the proprioception vector at the corresponding chunk and provides a derived gripper hint (OPEN, PARTIALLY_CLOSED, or CLOSED based on the inter-finger joint gap). ... The Gripper slot agrees with the proprioception reading in 82.7% of chunks overall."
The Gripper slot is supervised by a label derived from the proprioception vector, and the same proprioception reading is later used as the ground truth in the agreement check. The student also receives the proprioception vector as an input at inference, so high agreement measures whether the model reproduces a deterministic function of one of its inputs, not whether it independently tracks physical state from vision. This is a training-target sanity check, not an external validation, and it does not support the headline benchmark margins.
full rationale
The paper's central claim - that a 0.9B VLA exceeds 7B baselines on LIBERO-Plus - is an empirical benchmark comparison against externally published baseline numbers, not a derivation from fitted parameters. The reported margins (90.8/87.3/86.6/80.7) come from rollouts on the benchmark's perturbed tasks, and the ablations and interventions are matched within a single trained model. No load-bearing self-citation or uniqueness theorem is invoked; the method components (dual-view history, CoT distillation, paraphrase augmentation) are motivated by known perturbation axes but their effects on the benchmark are measured, not assumed. The Gripper-slot consistency check is self-referential because the validation signal is the same proprioception signal used to make the training label, and the paraphrase robustness probe is in-distribution since it draws on the augmentation set used during training; however, neither supports the headline result, which stands on the external LIBERO-Plus comparison. The unresolved question of whether baselines were trained on the same data as CoTinyVLA is a protocol-matching concern, not a circularity in the derivation.
Axiom & Free-Parameter Ledger
free parameters (5)
- CoT loss weight beta =
0.1
- History length per view =
8 frames per view
- Paraphrase swap probability =
0.8
- Reasoning token cap =
120
- Paraphrases per base instruction =
20
axioms (4)
- standard math LIBERO-Plus single-rollout evaluation, aggregated over ~2,500 tasks per suite, supports the reported binomial/independent-normal confidence intervals.
- domain assumption Training on the union of the four LIBERO-Plus suites and evaluating on LIBERO-Plus is a fair comparison; the published baselines had the same training-data access.
- domain assumption The 35B teacher's Plan/Think labels are sufficiently correct, and the proprioception-derived gripper hint is a valid training signal.
- domain assumption Paraphrase variants preserve the task semantics of the base instructions.
read the original abstract
Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus robustness benchmark use three- to seven-billion-parameter backbones whose memory demands can exceed embedded robotic budgets. We present CoTinyVLA, a 0.9B-parameter action model on a Qwen3.5-0.8B backbone that obtains that robustness by structuring supervision instead of enlarging the model. Three components target different axes of the problem: dual-view temporal input of 16 history frames per step with textual camera and time markers; hierarchical chain-of-thought (CoT) distillation from a 35B teacher into an episode-level Plan and a chunk-level Think span over task phase, gripper state and next subaction; and paraphrase augmentation expanding 40 base commands into 800 variants. On LIBERO-Plus, spanning 10,030 perturbed tasks across seven perturbation dimensions, CoTinyVLA reaches 90.8% on Spatial, 87.3% on Object, 86.6% on Goal and 80.7% on Long, leading the strongest 7B baseline on all four suites by 4.7, 2.8, 15.9 and 3.0 points, with every margin interval excluding zero. The gains concentrate on the hardest axes of the benchmark: across the eleven published baselines none exceeds 53.2% on Robot Initial States in any suite, whereas CoTinyVLA reaches 73.6% on Goal against 39.9% for the strongest baseline. Ablations show the three components to be separable by perturbation axis, and at a matched image budget how frames are divided between the two cameras and across time accounts for 8.6 points on its own. Closed-loop inference peaks at 2.25 GiB of allocated GPU memory, and paired interventions show the episode Plan to be load-bearing: replacing it with an empty or contradictory span costs 40 to 45 points of success. Structured supervision thus lets a 0.9B backbone exceed all of them. Code: https://github.com/BrainJellyPie/CoTinyVLA
Figures
Reference graph
Works this paper leans on
-
[1]
and Sanketi, Pannag R
Kim, Moo Jin and Pertsch, Karl and Karamcheti, Siddharth and Xiao, Ted and Balakrishna, Ashwin and Nair, Suraj and Rafailov, Rafael and Foster, Ethan P. and Sanketi, Pannag R. and Vuong, Quan and Kollar, Thomas and Burchfiel, Benjamin and Tedrake, Russ and Sadigh, Dorsa and Levine, Sergey and Liang, Percy and Finn, Chelsea , booktitle =. 2025 , editor =
2025
-
[2]
Proceedings of Robotics: Science and Systems , year =
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success , author =. Proceedings of Robotics: Science and Systems , year =
-
[3]
2025 , address =
Black, Kevin and Brown, Noah and Driess, Danny and Esmail, Adnan and Equi, Michael Robert and Finn, Chelsea and Fusai, Niccolo and Groom, Lachy and Hausman, Karol and Ichter, Brian and Jakubczak, Szymon and Jones, Tim and Ke, Liyiming and Levine, Sergey and Li-Bell, Adrian and Mothukuri, Mohith and Nair, Suraj and Pertsch, Karl and Shi, Lucy Xiaoyang and ...
2025
-
[4]
2025 , address =
Pertsch, Karl and Stachowicz, Kyle and Ichter, Brian and Driess, Danny and Nair, Suraj and Vuong, Quan and Mees, Oier and Finn, Chelsea and Levine, Sergey , booktitle =. 2025 , address =
2025
-
[5]
2025 , eprint =
Hung, Chia-Yu and Sun, Qi and Hong, Pengfei and Zadeh, Amir and Li, Chuan and Tan, U-Xuan and Majumder, Navonil and Poria, Soujanya , journal =. 2025 , eprint =
2025
-
[6]
2025 , eprint =
Cen, Jun and Yu, Chaohui and Yuan, Hangjie and Jiang, Yuming and Huang, Siteng and Guo, Jiayan and Li, Xin and Song, Yibing and Luo, Hao and Wang, Fan and Zhao, Deli and Chen, Hao , journal =. 2025 , eprint =
2025
-
[7]
The Fourteenth International Conference on Learning Representations , year =
Unified Vision-Language-Action Model , author =. The Fourteenth International Conference on Learning Representations , year =. 2506.19850 , archivePrefix =
-
[8]
Proceedings of Robotics: Science and Systems , year =
Learning to Act Anywhere with Task-centric Latent Actions , author =. Proceedings of Robotics: Science and Systems , year =
-
[9]
arXiv preprint arXiv:2505.17016 , year =
Interactive Post-Training for Vision-Language-Action Models , author =. arXiv preprint arXiv:2505.17016 , year =. 2505.17016 , archivePrefix =
-
[10]
Proceedings of the AAAI Conference on Artificial Intelligence , volume =
Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2026 , doi =
2026
-
[11]
2025 , eprint =
Shukor, Mustafa and Aubakirova, Dana and Capuano, Francesco and Kooijmans, Pepijn and Palma, Steven and Zouitine, Adil and Aractingi, Michel and Pascal, Caroline and Russi, Martino and Marafioti, Andres and Alibert, Simon and Cord, Matthieu and Wolf, Thomas and Cadene, Remi , journal =. 2025 , eprint =
2025
-
[12]
2024 , eprint =
Wen, Junjie and Zhu, Yichen and Li, Jinming and Zhu, Minjie and Wu, Kun and Xu, Zhiyuan and Liu, Ning and Cheng, Ran and Shen, Chaomin and Peng, Yaxin and Feng, Feifei and Tang, Jian , journal =. 2024 , eprint =
2024
-
[13]
2025 , address =
Qu, Delin and Song, Haoming and Chen, Qizhi and Yao, Yuanqi and Ye, Xinyi and Gu, Jiayuan and Wang, Zhigang and Ding, Yan and Zhao, Bin and Wang, Dong and Li, Xuelong , booktitle =. 2025 , address =
2025
-
[14]
Liu, Jiaming and Liu, Mengzhen and Wang, Zhenyu and An, Pengju and Li, Xiaoqi and Zhou, Kaichen and Yang, Senqiao and Zhang, Renrui and Guo, Yandong and Zhang, Shanghang , booktitle =. 2024 , doi =. 2406.04339 , archivePrefix =
Pith/arXiv arXiv 2024
-
[15]
2024 , eprint =
Li, Qixiu and Liang, Yaobo and Wang, Zeyu and Luo, Lin and Chen, Xi and Liao, Mozheng and Wei, Fangyun and Deng, Yu and Xu, Sicheng and Zhang, Yizhong and Wang, Xiaofan and Liu, Bei and Fu, Jianlong and Bao, Jianmin and Chen, Dong and Shi, Yuanchun and Yang, Jiaolong and Guo, Baining , journal =. 2024 , eprint =
2024
-
[16]
Liu, Songming and Wu, Lingxuan and Li, Bangguo and Tan, Hengkai and Chen, Huayu and Wang, Zhengyi and Xu, Ke and Su, Hang and Zhu, Jun , booktitle =
-
[17]
and Salazar, Grecia and Sanketi, Pannag R
Brohan, Anthony and Brown, Noah and Carbajal, Justice and Chebotar, Yevgen and Dabis, Joseph and Finn, Chelsea and Gopalakrishnan, Keerthana and Hausman, Karol and Herzog, Alexander and Hsu, Jasmine and Ibarz, Julian and Ichter, Brian and Irpan, Alex and Jackson, Tomas and Jesmonth, Sally and Joshi, Nikhil and Julian, Ryan and Kalashnikov, Dmitry and Kuan...
2023
-
[18]
and Salazar, Grecia and Ryoo, Michael S
Zitkovich, Brianna and Yu, Tianhe and Xu, Sichun and Xu, Peng and Xiao, Ted and Xia, Fei and Wu, Jialin and Wohlhart, Paul and Welker, Stefan and Wahid, Ayzaan and Vuong, Quan and Vanhoucke, Vincent and Tran, Huong and Soricut, Radu and Singh, Anikait and Singh, Jaspiar and Sermanet, Pierre and Sanketi, Pannag R. and Salazar, Grecia and Ryoo, Michael S. a...
2023
-
[19]
Driess, Danny and Xia, Fei and Sajjadi, Mehdi S. M. and Lynch, Corey and Chowdhery, Aakanksha and Ichter, Brian and Wahid, Ayzaan and Tompson, Jonathan and Vuong, Quan and Yu, Tianhe and Huang, Wenlong and Chebotar, Yevgen and Sermanet, Pierre and Duckworth, Daniel and Levine, Sergey and Vanhoucke, Vincent and Hausman, Karol and Toussaint, Marc and Greff,...
2023
-
[20]
Transactions on Machine Learning Research , year=
A Generalist Agent , author=. Transactions on Machine Learning Research , year=
-
[21]
2022 , editor=
Jang, Eric and Irpan, Alex and Khansari, Mohi and Kappler, Daniel and Ebert, Frederik and Lynch, Corey and Levine, Sergey and Finn, Chelsea , booktitle=. 2022 , editor=
2022
-
[22]
2022 , editor=
Shridhar, Mohit and Manuelli, Lucas and Fox, Dieter , booktitle=. 2022 , editor=
2022
-
[23]
Proceedings of The 6th Conference on Robot Learning , pages =
Perceiver-Actor: A Multi-Task Transformer for Robotic Manipulation , author =. Proceedings of The 6th Conference on Robot Learning , pages =. 2023 , editor =
2023
-
[24]
2023 , editor =
Jiang, Yunfan and Gupta, Agrim and Zhang, Zichen and Wang, Guanzhi and Dou, Yongqiang and Chen, Yanjun and Fei-Fei, Li and Anandkumar, Anima and Zhu, Yuke and Fan, Linxi , booktitle =. 2023 , editor =
2023
-
[25]
The Twelfth International Conference on Learning Representations , year =
Vision-Language Foundation Models as Effective Robot Imitators , author =. The Twelfth International Conference on Learning Representations , year =. 2311.01378 , archivePrefix =
-
[26]
and Sadigh, Dorsa and Finn, Chelsea and Levine, Sergey , booktitle =
Ghosh, Dibya and Walke, Homer Rich and Pertsch, Karl and Black, Kevin and Mees, Oier and Dasari, Sudeep and Hejna, Joey and Kreiman, Tobias and Xu, Charles and Luo, Jianlan and Tan, You Liang and Chen, Lawrence Yunliang and Vuong, Quan and Xiao, Ted and Sanketi, Pannag R. and Sadigh, Dorsa and Finn, Chelsea and Levine, Sergey , booktitle =. 2024 , address =
2024
-
[27]
2024 , doi =
IEEE International Conference on Robotics and Automation (ICRA) , pages =. 2024 , doi =
2024
-
[28]
Walke, Homer Rich and Black, Kevin and Zhao, Tony Z. and Vuong, Quan and Zheng, Chongyi and Hansen-Estruch, Philippe and He, Andre Wang and Myers, Vivek and Kim, Moo Jin and Du, Max and Lee, Abraham and Fang, Kuan and Finn, Chelsea and Levine, Sergey , booktitle=. 2023 , editor=
2023
-
[29]
Proceedings of Robotics: Science and Systems , year =
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion , author =. Proceedings of Robotics: Science and Systems , year =
-
[30]
Proceedings of Robotics: Science and Systems , year =
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware , author =. Proceedings of Robotics: Science and Systems , year =
-
[31]
The Twelfth International Conference on Learning Representations , year =
Gu, Jiayuan and Kirmani, Sean and Wohlhart, Paul and Lu, Yao and. The Twelfth International Conference on Learning Representations , year =
-
[32]
Zhao, Qingqing and Lu, Yao and Kim, Moo Jin and Fu, Zipeng and Zhang, Zhuoyang and Wu, Yecheng and Li, Zhaoshuo and Ma, Qianli and Han, Song and Finn, Chelsea and Handa, Ankur and Lin, Tsung-Yi and Wetzstein, Gordon and Liu, Ming-Yu and Xiang, Donglai , booktitle =
-
[33]
Proceedings of The 8th Conference on Robot Learning , pages =
Robotic Control via Embodied Chain-of-Thought Reasoning , author =. Proceedings of The 8th Conference on Robot Learning , pages =. 2025 , volume =
2025
-
[34]
2026 , eprint =
Zhong, Linqing and Liu, Yi and Wei, Yifei and Xiong, Ziyu and Yao, Maoqing and Liu, Si and Ren, Guanghui , journal =. 2026 , eprint =
2026
-
[35]
Liu, Bo and Zhu, Yifeng and Gao, Chongkai and Feng, Yihao and Liu, Qiang and Zhu, Yuke and Stone, Peter , booktitle=
-
[36]
2025 , eprint =
Fei, Senyu and Wang, Siyin and Shi, Junhao and Dai, Zihao and Cai, Jikun and Qian, Pengfang and Ji, Li and He, Xinzhe and Zhang, Shiduo and Fei, Zhaoye and Fu, Jinlan and Gong, Jingjing and Qiu, Xipeng , journal =. 2025 , eprint =
2025
-
[37]
, journal=
James, Stephen and Ma, Zicong and Arrojo, David Rovick and Davison, Andrew J. , journal=. 2020 , doi=
2020
-
[38]
2022 , doi =
Mees, Oier and Hermann, Lukas and Rosete-Beas, Erick and Burgard, Wolfram , journal =. 2022 , doi =
2022
-
[39]
2020 , volume =
Yu, Tianhe and Quillen, Deirdre and He, Zhanpeng and Julian, Ryan and Hausman, Karol and Finn, Chelsea and Levine, Sergey , booktitle =. 2020 , volume =
2020
-
[40]
arXiv preprint arXiv:2009.12293 , year =
Zhu, Yuke and Wong, Josiah and Mandlekar, Ajay and Mart. arXiv preprint arXiv:2009.12293 , year =. 2009.12293 , archivePrefix =
Pith/arXiv arXiv 2009
-
[41]
Proceedings of the 5th Conference on Robot Learning , pages=
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation , author=. Proceedings of the 5th Conference on Robot Learning , pages=. 2022 , editor=
2022
-
[42]
2023 , editor=
Mandlekar, Ajay and Nasiriany, Soroush and Wen, Bowen and Akinola, Iretiayo and Narang, Yashraj and Fan, Linxi and Zhu, Yuke and Fox, Dieter , booktitle=. 2023 , editor=
2023
-
[43]
Zhang, Shiduo and Xu, Zhe and Liu, Peiju and Yu, Xiaopeng and Li, Yuan and Gao, Qinghui and Fei, Zhaoye and Yin, Zhangyue and Wu, Zuxuan and Jiang, Yu-Gang and Qiu, Xipeng , booktitle =
-
[44]
Advances in Neural Information Processing Systems , volume=
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models , author=. Advances in Neural Information Processing Systems , volume=
-
[45]
Findings of the Association for Computational Linguistics: ACL 2023 , pages=
Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes , author=. Findings of the Association for Computational Linguistics: ACL 2023 , pages=. 2023 , address=
2023
-
[46]
Transactions on Machine Learning Research , year=
Multimodal Chain-of-Thought Reasoning in Language Models , author=. Transactions on Machine Learning Research , year=. 2302.00923 , archivePrefix=
-
[47]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Large Language Models Are Reasoning Teachers , author=. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2023 , address=
2023
-
[48]
2023 , doi=
Bai, Jinze and Bai, Shuai and Yang, Shusheng and Wang, Shijie and Tan, Sinan and Wang, Peng and Lin, Junyang and Zhou, Chang and Zhou, Jingren , journal=. 2023 , doi=
2023
-
[49]
2025 , eprint =
Bai, Shuai and Chen, Keqin and Liu, Xuejing and Wang, Jialin and Ge, Wenbin and Song, Sibo and Dang, Kai and Wang, Peng and Wang, Shijie and Tang, Jun and Zhong, Humen and Zhu, Yuanzhi and Yang, Mingkun and Li, Zhaohai and Wan, Jianqiang and Wang, Pengfei and Ding, Wei and Fu, Zheren and Xu, Yiheng and Ye, Jiabo and Zhang, Xi and Xie, Tianbao and Cheng, Z...
2025
-
[50]
2026 , howpublished =
2026
-
[51]
Advances in Neural Information Processing Systems , volume=
Visual Instruction Tuning , author=. Advances in Neural Information Processing Systems , volume=
-
[52]
2023 , editor=
Li, Junnan and Li, Dongxu and Savarese, Silvio and Hoi, Steven , booktitle=. 2023 , editor=
2023
-
[53]
Transactions on Machine Learning Research , year =
Oquab, Maxime and Darcet, Timoth. Transactions on Machine Learning Research , year =. 2304.07193 , archivePrefix =
-
[54]
Zhai, Xiaohua and Mustafa, Basil and Kolesnikov, Alexander and Beyer, Lucas , booktitle=
-
[55]
2023 , doi=
Awadalla, Anas and Gao, Irena and Gardner, Josh and Hessel, Jack and Hanafy, Yusuf and Zhu, Wanrong and Marathe, Kalyani and Bitton, Yonatan and Gadre, Samir and Sagawa, Shiori and Jitsev, Jenia and Kornblith, Simon and Koh, Pang Wei and Ilharco, Gabriel and Wortsman, Mitchell and Schmidt, Ludwig , journal=. 2023 , doi=
2023
-
[56]
2019 , address=
Wei, Jason and Zou, Kai , booktitle=. 2019 , address=
2019
-
[57]
Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Improving Neural Machine Translation Models with Monolingual Data , author=. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=. 2016 , address=
2016
-
[58]
Edward and Rudinger, Rachel and Post, Matt and Van Durme, Benjamin , booktitle=
Hu, J. Edward and Rudinger, Rachel and Post, Matt and Van Durme, Benjamin , booktitle=. 2019 , doi=
2019
-
[59]
2021 , address=
Karimi, Akbar and Rossi, Leonardo and Prati, Andrea , booktitle=. 2021 , address=
2021
-
[60]
2026 , eprint =
Zhang, Yichi and Yuan, Weihao and Zhang, Yizhuo and Zhang, Xidong and Wan, Jia , journal =. 2026 , eprint =
2026
-
[61]
2026 , eprint =
Luo, Yuankai and Chen, Woping and Liang, Tong and Wang, Baiqiao and Li, Zhenguo , journal =. 2026 , eprint =
2026
-
[62]
arXiv preprint arXiv:2602.20200 , year =
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation , author =. arXiv preprint arXiv:2602.20200 , year =. 2602.20200 , archivePrefix =
-
[63]
Wang, Yihao and Ding, Pengxiang and Li, Lingxiao and Cui, Can and Ge, Zirui and Tong, Xinyang and Song, Wenxuan and Zhao, Han and Zhao, Wei and Hou, Pengxu and Huang, Siteng and Tang, Yifan and Wang, Wenhui and Zhang, Ru and Liu, Jianyi and Wang, Donglin , journal =. 2026 , doi =. 2509.09372 , archivePrefix =
arXiv 2026
-
[64]
2025 , eprint =
Zheng, Jinliang and Li, Jianxiong and Wang, Zhihao and Liu, Dongxiu and Kang, Xirui and Feng, Yuchun and Zheng, Yinan and Zou, Jiayin and Chen, Yilun and Zeng, Jia and Zhang, Ya-Qin and Pang, Jiangmiao and Liu, Jingjing and Wang, Tai and Zhan, Xianyuan , journal =. 2025 , eprint =
2025
-
[65]
Black, Kevin and Brown, Noah and Darpinian, James and Dhabalia, Karan and Driess, Danny and Esmail, Adnan and Equi, Michael Robert and Finn, Chelsea and Fusai, Niccolo and Galliker, Manuel Y. and Ghosh, Dibya and Groom, Lachy and Hausman, Karol and Ichter, Brian and Jakubczak, Szymon and Jones, Tim and Ke, Liyiming and LeBlanc, Devin and Levine, Sergey an...
2025
-
[66]
Proceedings of The 9th Conference on Robot Learning , pages =
Reuss, Moritz and Zhou, Hongyi and R. Proceedings of The 9th Conference on Robot Learning , pages =. 2025 , editor =
2025
-
[67]
2025 , eprint =
Ni, Chaojun and Chen, Cheng and Wang, Xiaofeng and Zhu, Zheng and Zheng, Wenzhao and Wang, Boyuan and Chen, Tianrun and Zhao, Guosheng and Li, Haoyun and Dong, Zhehao and Zhang, Qiang and Ye, Yun and Wang, Yang and Huang, Guan and Mei, Wenjun , journal =. 2025 , eprint =
2025
-
[68]
2025 , eprint =
Hung, Chia-Yu and Majumder, Navonil and Deng, Haoyuan and Liu, Renhang and Ang, Yankang and Zadeh, Amir and Li, Chuan and Herremans, Dorien and Wang, Ziwei and Poria, Soujanya , journal =. 2025 , eprint =
2025
-
[69]
2025 , eprint =
Li, Yixuan and Chen, Yuhui and Zhou, Mingcai and Li, Haoran and Zhang, Zhengtao and Zhao, Dongbin , journal =. 2025 , eprint =
2025
-
[70]
2025 , eprint =
Lin, Tao and Zhong, Yilei and Du, Yuxin and Zhang, Jingjing and Liu, Jiting and Chen, Yinxinyu and Gu, Encheng and Liu, Ziyan and Cai, Hongyi and Zou, Yanwen and Zou, Lixing and Zhou, Zhaoye and Li, Gen and Zhao, Bo , journal =. 2025 , eprint =
2025
-
[71]
2025 , eprint =
Goyal, Ankit and Hadfield, Hugo and Yang, Xuning and Blukis, Valts and Ramos, Fabio , journal =. 2025 , eprint =
2025
-
[72]
Balakhnov, Oleg and Skvortsov, Sergei and Zarin, Gleb , journal =
-
[73]
arXiv preprint arXiv:2503.14734 , year =. 2503.14734 , archivePrefix =
-
[74]
Recurrent-Depth
Tur, Yalcin and Naghiyev, Jalal and Fang, Haoquan and Tsai, Wei-Chuan and Duan, Jiafei and Fox, Dieter and Krishna, Ranjay , journal =. Recurrent-Depth. 2026 , eprint =
2026
-
[75]
2026 , eprint =
Li, Dachong and Chen, ZhuangZhuang and Zhang, Jin and Li, Jianqiang , journal =. 2026 , eprint =
2026
-
[76]
2025 , eprint =
Deng, Shengliang and Yan, Mi and Wei, Songlin and Ma, Haixin and Yang, Yuxin and Chen, Jiayi and Zhang, Zhiqi and Yang, Taoyu and Zhang, Xuheng and Cui, Heming and Zhang, Zhizheng and Wang, He , journal =. 2025 , eprint =
2025
-
[77]
2025 , eprint =
Xiong, Zheng and Li, Kang and Wang, Zilin and Jackson, Matthew and Foerster, Jakob and Whiteson, Shimon , journal =. 2025 , eprint =
2025
-
[78]
2025 , eprint =
Zhang, Jiahui and Chen, Yurui and Xu, Yueming and Huang, Ze and Zhou, Yanpeng and Yuan, Yu-Jie and Cai, Xinyue and Huang, Guowei and Quan, Xingyue and Xu, Hang and Zhang, Li , journal =. 2025 , eprint =
2025
-
[79]
2025 , eprint =
Zhang, Ji and Wu, Shihan and Luo, Xu and Wu, Hao and Gao, Lianli and Shen, Heng Tao and Song, Jingkuan , journal =. 2025 , eprint =
2025
-
[80]
2025 , eprint =
Gao, Chongkai and Liu, Zixuan and Chi, Zhenghao and Huang, Junshan and Fei, Xin and Hou, Yiwen and Zhang, Yuxuan and Lin, Yudi and Fang, Zhirui and Jiang, Zeyu and Shao, Lin , journal =. 2025 , eprint =
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.