Pith. sign in

REVIEW 2 major objections 4 minor 60 references

RADIO1D: Elastic Representations for Condensed Vision Modeling

T0 review · 2 major / 4 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read VLMs do not need fixed 2D patch grids; variable-length 1D tokens can match or beat them while letting you dial accuracy against cost.

desk verdict Solid systems paper: elastic continuous 1D tokens from multi-teacher distillation give a real VLM accuracy-latency Pareto front and better composition retrieval, with the hierarchical-ordering story only partially isolated. read the letter →

arxiv 2607.03624 v1 pith:5PAEBXSG submitted 2026-07-03 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords vision-languagemodels1Dtokenizationelasticrepresentationsmulti-teacherdistillationnesteddropouthierarchicalsummarizationcomposition-awareretrievaltokenefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Vision-language models have long assumed that the vision side must hand the language model a fixed grid of 2D patch tokens. This paper shows that assumption is unnecessary. When vision encoders are fine-tuned inside VLMs, their features become more abstract and less spatially coherent; image-text models already invent a few specialized tokens that act as global summaries. RADIO1D builds on that observation by training a student that compresses any image into a short, ordered 1D sequence whose length can be chosen at inference time. Nested dropout forces the earliest tokens to carry the most global information, so even a single token already supports useful scene understanding and better composition-aware retrieval. When plugged into a 9B language model, the same checkpoint yields a continuous accuracy-latency curve that is competitive with or better than fixed-256-token baselines while using far fewer tokens on many tasks.

What carries the argument

RADIO1D: an encoder-decoder student that maps an image to an elastic 1D token sequence via multi-teacher distillation (SigLIP2, DINOv3, SAM3) and nested dropout (triangular length prior, early tokens kept preferentially), with Patch Merging internalized at block 24; the decoder is used only at training time to restore a 2D grid for teacher alignment.

What would settle it

Train an otherwise identical continuous 1D auto-encoder without nested dropout (or with a uniform length prior) and measure whether the single-token and low-rate ADE20K mIoU and VLM scores collapse relative to the nested-dropout RADIO1D checkpoint.

Watch

Extended reading notes

Core claim

Fixed 2D patch grids are not required for strong VLM performance. A hierarchical 1D sequence produced by multi-teacher distillation and nested dropout can summarize an image so effectively that a single token already yields non-trivial scene understanding, and adjustable token counts give continuous accuracy-efficiency trade-offs that match or exceed fixed-grid baselines on multimodal benchmarks.

Load-bearing premise

The paper treats the hierarchical ordering created by nested dropout plus the chosen multi-teacher mix as the main reason for strong low-token performance, without fully isolating that mechanism from the continuous embedding space or the rest of the auto-encoder design.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper argues that VLMs do not require fixed 2D patch grids. Analysis of SigLIP2, DINOv3 and VLM-finetuned C-RADIOv4 shows that image–text training produces abstract, low-spatial-coherence features and a few specialized global tokens. RADIO1D is then introduced: a multi-teacher (SigLIP2, DINOv3, SAM3) distilled encoder that maps an image to a continuous, variable-length 1D sequence via nested dropout (triangular length prior) and an internal Patch-Merging stage (κ=24, ρ=2). A training-only decoder reconstructs 2D features for teacher alignment. Experiments demonstrate hierarchical summarization (single-token ADE20K mIoU 40.23, strong ImageNet k-NN on early tokens), superior rate–distortion curves, a new composition-aware retrieval metric (Comp@K) on which RADIO1D outperforms baselines, and a continuous accuracy–latency Pareto front on ten VLM benchmarks when the same checkpoint is sliced to 1–256 tokens and paired with a fixed 9B Nemotron LLM.

Significance. If the empirical results hold, the work supplies a practical, elastic vision backbone that lets practitioners trade accuracy for latency at inference time without retraining, while matching or exceeding fixed-token SigLIP2 and C-RADIOv4 baselines. The composition-aware retrieval metric and the controlled Nemotron-VL ablation (identical LLM, data and schedule) are concrete, reusable contributions. The analysis of specialized global tokens and the loss of spatial coherence under VLM fine-tuning is also of independent interest to the foundation-model community. Model checkpoints are released under a permissive license, increasing the work’s immediate utility.

major comments (2)
  1. Section 4.1 and Figure 5 ablate the length prior and down-scaling position, yet never isolate nested dropout itself from the rest of the auto-encoder + multi-teacher recipe. Because hierarchical ordering is repeatedly invoked as the mechanism behind the low-rate regime (Figures 8–9, single-token mIoU 40.23), a controlled ablation that freezes the architecture and teachers while removing or randomizing the nested-dropout schedule would strengthen the causal claim; without it the systems-level Pareto front remains solid but the mechanistic story is only correlational.
  2. Table 1 reports single-run VLM numbers with no error bars or multi-seed statistics. Given that the central claim is a continuous accuracy–latency trade-off that “matches or exceeds” fixed-256-token baselines, modest run-to-run variance could reorder the Pareto ranking at intermediate token counts (e.g., 128 vs. 192). At least three independent SFT seeds for the key RADIO1D points would make the comparison statistically robust.
minor comments (4)
  1. Figure 1 caption and Table 1: TTFT is measured with a fixed 32-image / 128-token context; a short note on how the measurement changes with variable tile counts would help readers extrapolate to other deployment settings.
  2. Section 3.3: the claim “Minimizing the objective ensures that I(T_i;Y)>I(T_j;Y) ∀i<j” is stated without a formal derivation; a short information-theoretic argument or a pointer to the nested-dropout literature would clarify the guarantee.
  3. Appendix L (CKA regularization) is interesting but ultimately discarded; a one-sentence summary in the main text of why the regularizer was abandoned would prevent readers from wondering whether the final model still contains residual 2-D structure.
  4. Typographical: “A verage” in Table 1 header; “L VBench” should be “LVBench”; “pre-trained C-RADIOv4-H and C-RADIOv4-H (fine-tuned in a VLM)” in Figure 6 caption is slightly redundant.

Circularity Check

1 steps flagged · score 1.0 of 10

No significant circularity: empirical VLM and retrieval results rest on external benchmarks and standard multi-teacher distillation; self-citations supply only the base 2D RADIO recipe.

  1. self citation load bearing [Section 3.1 Agglomerative Training; also D. Initialization]
    "We adopt the training recipe of C-RADIOv4 [26] to produce RADIO1D. ... We initialize RADIO1D from a pre-trained standard 2D Vision Transformer checkpoint (e.g., C-RADIOv4) using a dedicated conversion procedure."

    The base agglomerative multi-teacher recipe and weight-initialization procedure are taken from the authors’ own prior C-RADIO papers. This is ordinary engineering reuse, not a load-bearing uniqueness claim or a definition that forces the new 1D elastic numbers; the variable-length results and VLM tables remain independent measurements.

full rationale

RADIO1D is an empirical systems paper. Its load-bearing claims (variable-token VLM accuracy–latency Pareto front on ten public benchmarks, single-token ADE20K mIoU, composition-aware retrieval on MS-COCO/nuImages) are measured against held-out external data with a controlled Nemotron-VL setup that freezes the LLM, data, and schedule. Nested dropout and the triangular length prior are training regularizers whose hierarchical effect is verified post-hoc by prefix/suffix curves and rate-distortion plots, not definitional identities that force the reported numbers. Multi-teacher losses are ordinary MSE/cosine distillation. The sole self-citations (C-RADIOv4 training recipe, Net2WiderNet-style initialization) supply the starting 2D backbone and are not used to prove uniqueness or to derive the 1D elastic results; those results are new measurements. No fitted parameter is renamed a prediction, no uniqueness theorem is imported, and no known result is merely re-labeled. Score 1 reflects only the minor, non-load-bearing self-citation of the authors’ prior RADIO line.

Assumptions & free parameters 4 free parameters · 3 assumptions · 2 invented entities

The central empirical claims rest on a handful of design choices (triangular length prior, Patch-Merging location κ=24, expansion ρ=2, three specific teachers) that were selected by ablation on ADE20K and then frozen. No free parameters are fitted to the final VLM numbers; the composition metric introduces a new scoring function whose validity is assumed rather than derived.

free parameters (4)
  • triangular PDF for token length ℓ = p(x)=2-2x on [0,1]
    Chosen after ablation (Fig. 5) because it concentrates mass on short sequences; the exact functional form 2-2x is hand-selected.
  • downscaling block position κ = 24
    Selected as κ=24 after short-schedule ADE20K sweeps; earlier or later positions degrade mIoU.
  • embedding expansion factor ρ = 2
    Fixed to 2 after static throughput/parameter analysis; higher values explode parameter count.
  • teacher loss weights λ_t
    Inherited from C-RADIOv4 recipe; not re-tuned for the 1D student.
assumptions (3)
  • domain assumption Nested dropout on a 1D sequence induces a strict hierarchical ordering of mutual information I(T_i;Y) > I(T_j;Y) for i<j
    Stated in §3.3 and used to justify the elastic design; empirically supported but not proved.
  • domain assumption Multi-teacher continuous distillation from DINOv3, SigLIP2 and SAM3 yields a student whose early tokens are globally informative
    Core training objective of §3.1; success is measured post-hoc rather than guaranteed by theory.
  • ad hoc to paper The bipartite gIoU-based composition score correctly quantifies scene layout similarity for retrieval
    Introduced in §4.6; no external validation that higher Comp@K predicts human judgments of composition.
invented entities (2)
  • RADIO1D elastic 1D continuous token sequence independent evidence
    purpose: Replace fixed 2D patch grids with a variable-length hierarchical summary usable by VLMs
    The core architectural contribution; independent evidence is the released checkpoints and the reported benchmarks.
  • Composition-aware retrieval metric (Comp@K)
    purpose: Score whether retrieved images share object categories, sizes and spatial layout with the query
    New evaluation protocol defined via Hungarian matching of gIoU; no prior literature uses this exact score.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RADIO1D: Elastic Representations for Condensed Vision Modeling." pith.science (2026). https://pith.science/paper/5PAEBXSG

@misc{pith2026260703624,
  author       = {Pith},
  title        = {Pith review of: RADIO1D: Elastic Representations for Condensed Vision Modeling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5PAEBXSG}},
  note         = {Machine review of arXiv:2607.03624}
}
read the original abstract

This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned vision encoders, we find that representations become increasingly abstract and less spatially coherent during VLM training. Notably, models trained with image-text alignment (such as SigLIP2) develop a small number of specialized tokens that effectively summarize global image content. Building on this, we introduce RADIO1D, which compresses images into a compact, variable-length 1D token sequence using multi-teacher knowledge distillation and an autoencoder design. The resulting representations exhibit strong hierarchical summarization, enabling accurate scene understanding - even with a single token - and support improved composition-aware image retrieval. In VLMs, RADIO1D provides flexible accuracy-efficiency tradeoffs through adjustable token counts, delivering competitive performance on diverse multimodal benchmarks with lower computational overhead and better accuracy.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

60 extracted references · 2 canonical work pages

  1. [1]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang, editors,Proceedings of the 38th International Conference on Machin...

  2. [2]

    BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models. In Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara En- gelhardt, Sivan Sabato, and Jonathan Scarlett, editors,Proceedings of the 40th International Conference on Machine Learning, volume 202...

  3. [3]

    Language is not all you need: Aligning perception with language models

    Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao, Saksham Singhal, Shuming Ma, Tengchao Lv, Lei Cui, Owais Khan Mohammed, Barun Patra, Qiang Liu, Kriti Aggarwal, Zewen Chi, Johan Bjorck, Vishrav Chaudhary, Subhojit Som, Xia Song, and Furu Wei. Language is not all you need: Aligning perception with language models. InAdvances in Neural Information Processing ...

  4. [4]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. InAd- vances in Neural Information Processing Systems, volume 36, 2023

  5. [5]

    Sigmoid loss for language image pre-training

    Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 11975–11986, October 2023

  6. [6]

    Siglip 2: Multilingual vision- language encoders with improved semantic under- standing, localization, and dense features, 2025

    Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alab- dulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hé- naff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. Siglip 2: Multilingual vision- language encoders with improved semantic under- standing, localization, and dense featu...

  7. [7]

    Paligemma: A versatile 3b vlm for transfer, 2024

    Lucas Beyer, Andreas Steiner, André Su- sano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Al- abdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias...

  8. [8]

    What matters when building vision-language models?, 2024

    Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. What matters when building vision-language models?, 2024. URLhttps:// arxiv.org/abs/2405.02246

Show all 60 references
  1. [9]

    Cambrian-1: A fully open, vision-centric exploration of multimodal llms

    Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann Le- Cun, and Saining Xie. Cambrian-1: A fully open, vision-centric exploration of multimodal llms. In ...

  2. [10]

    Nvila: Efficient frontier visual language models, 2024

    Zhijian Liu, Ligeng Zhu, Baifeng Shi, Zhuoyang Zhang, Yuming Lou, Shang Yang, Haocheng Xi, Shiyi Cao, Yuxian Gu, Dacheng Li, Xiuyu Li, Yunhao Fang, Yukang Chen, Cheng-Yu Hsieh, De-An Huang, An-Chieh Cheng, Vishwesh Nath, Jinyi Hu, Sifei Liu, Ranjay Krishna, Daguang Xu, Xiaolon...

  3. [11]

    Qwen3-vl technical report, 2025

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junya...

  4. [12]

    Radiov2.5: Improved baselines for agglomerative vision foun- dation models

    Greg Heinrich, Mike Ranzinger, Hongxu Danny Yin, Yao Lu, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov. Radiov2.5: Improved baselines for agglomerative vision foun- dation models. In2025 IEEE/CVF Confer- ence on Computer Vision and Pattern Recog- nition (CVPR), p...

  5. [13]

    Amala Sanjay Deshmukh, Kateryna Chu- machenko, Tuomas Rintamaki, Matthieu Le, Tyler Poon, Danial Mohseni-Taheri, Ilia Kar- manov, Guilin Liu, Jarno Seppänen, Guo Chen, Karan Sapra, Zhi-Wei Yu, Adi Renduchin- tala, Charles Wang, Peter Jin, Arushi Goel, Mike Ranzinger, Lukas Voe...

  6. [14]

    Oriane Siméoni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Francisco Massa, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Tim- othée Darcet, Théo Moutakanni, Leonel Sen- tana...

  7. [15]

    Sam 3: Segment anything with 11 RADIO1D: Elastic Representations for Condensed Vision Modeling concepts, 2025

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, and et al. Sam 3: Segment anything with 11 RADIO1D: Elastic Representations for Condensed Vision Modeling concepts, 2025. URL https://arxiv.org/abs/ 2511.16719

  8. [16]

    Eagle: Exploring the design space for multimodal LLMs with mixture of encoders

    Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, Yilin Zhao, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, and Guilin Liu. Eagle: Exploring the design space for multimodal LLMs with...

  9. [17]

    VILA-u: a unified foundation model integrating visual understanding and generation

    Yecheng Wu, Zhuoyang Zhang, Junyu Chen, Hao- tian Tang, Dacheng Li, Yunhao Fang, Ligeng Zhu, Enze Xie, Hongxu Yin, Li Yi, Song Han, and Yao Lu. VILA-u: a unified foundation model integrating visual understanding and generation. InThe Thirteenth International Conference on Lear...

  10. [18]

    Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,

  11. [19]

    URL https://openreview.net/forum? id=qrGjFJVl3m

  12. [20]

    Llava- uhd v2: an mllm integrating high-resolution se- mantic pyramid via hierarchical window trans- former, 2025

    Yipeng Zhang, Yifan Liu, Zonghao Guo, Yi- dan Zhang, Xuesong Yang, Xiaoying Zhang, Chi Chen, Jun Song, Bo Zheng, Yuan Yao, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun. Llava- uhd v2: an mllm integrating high-resolution se- mantic pyramid via hierarchical window trans- former, ...

  13. [21]

    Scene parsing through ADE20K dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene parsing through ADE20K dataset. InPro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5122–5130, 2017. doi: 10.1109/CVPR.2017.544. URL https:...

  14. [22]

    Schwing, Alexander Kirillov, and Rohit Girdhar

    Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for univer- sal image segmentation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1290–1299, 2022

  15. [23]

    Donghao Zhang, Yimin Chen, Kauê T. N. Duarte, Taha Aslan, Mohamed AlShamrani, Brij Karmur, Yan Wan, Shengcai Chen, Bo Hu, Bi- joy K. Menon, and Wu Qiu. Benchmarking dinov3 for multi-task stroke analysis on non- contrast ct, 2025. URL https://arxiv.org/ abs/2509.23132

  16. [24]

    DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026

    Zhiyi Dou, Edore Akpokodje, Yuelin He, Yuxin Liu, Zixuan Ni, Chang’an Xu, Muhammad Aslam, and Meng Tang. DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026. doi: 10.3390/s26020406. URL https://doi.org/10. 3390/s26020406

  17. [25]

    Similarity of neural network representations revisited

    Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network representations revisited. In Kamalika Chaudhuri and Ruslan Salakhutdi- nov, editors,Proceedings of the 36th Interna- tional Conference on Machine Learning, vol- ume 97 ofProceedi...

  18. [26]

    URL https://proceedings.mlr.press/ v97/kornblith19a.html

  19. [27]

    Lawrence Zit- nick, and Piotr Dollár

    Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zit- nick, and Piotr Dollár. Microsoft coco: Common objects in context, 2015

  20. [28]

    C-radiov4 (tech report), 2026

    Mike Ranzinger, Greg Heinrich, Collin McCarthy, Jan Kautz, Andrew Tao, Bryan Catanzaro, and Pavlo Molchanov. C-radiov4 (tech report), 2026. URLhttps://arxiv.org/abs/2601.17237

  21. [29]

    Pereira, and William Bialek

    Naftali Tishby, Fernando C. Pereira, and William Bialek. The information bottle- neck method. InProceedings of the 37th Annual Allerton Conference on Communica- tion, Control, and Computing, pages 368– 377, 1999. URL https://www.cs.huji.ac.il/ ~tishby/papers/IB-Allerton.pdf

  22. [30]

    Modeling by shortest data description.Automatica, 14(5):465–471, 1978

    Jorma Rissanen. Modeling by shortest data description.Automatica, 14(5):465–471, 1978. doi: 10.1016/0005-1098(78)90005-5. URLhttps: //doi.org/10.1016/0005-1098(78)90005-5

  23. [31]

    AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One

    Mike Ranzinger, Greg Heinrich, Jan Kautz, and Pavlo Molchanov. AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One . In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 12490–12500, Los Alamitos, CA, USA, June 2024. IEEE ...

  24. [32]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexan- der Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...

  25. [33]

    Flextok: Resam- pling images into 1d token sequences of flexible length

    Roman Bachmann, Jesse Allardice, David Mizrahi, Enrico Fini, Oğuzhan Fatih Kar, Elmira Amirloo, Alaaeldin El-Nouby, Amir Za- mir, and Afshin Dehghan. Flextok: Resam- pling images into 1d token sequences of flexible length. InForty-second International Confer- ence on Machine L...

  26. [34]

    Swin transformer: Hierarchical vision transformer using shifted windows

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yix- uan Wei, Zheng Zhang, Stephen Lin, and Bain- ing Guo. Swin transformer: Hierarchical vision transformer using shifted windows. InProceed- ings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 10012–10022, October 2021

  27. [35]

    Efros, Jenia Jitsev, Yair Carmon, Lud- wig Schmidt, and Vaishaal Shankar

    Samir Yitzhak Gadre, Gabriel Ilharco, Alex Fang, Jonathan Hayase, Georgios Smyrnis, Thao Nguyen, Ryan Marten, Mitchell Wortsman, Dhruba Ghosh, Jieyu Zhang, Ekin Dogus Cubuk, Alexei A. Efros, Jenia Jitsev, Yair Carmon, Lud- wig Schmidt, and Vaishaal Shankar. Datacomp: In search...

  28. [36]

    Getting vit in shape: Scaling laws for compute-optimal model design

    Ibrahim Alabdulmohsin, Xiaohua Zhai, Alexan- der Kolesnikov, and Lucas Beyer. Getting vit in shape: Scaling laws for compute-optimal model design. InAdvances in Neural Informa- tion Processing Systems, volume 36, 2023. URL https://arxiv.org/abs/2305.13035

  29. [37]

    Cover and Joy A

    Thomas M. Cover and Joy A. Thomas. Rate distortion theory. InElements of Informa- tion Theory, chapter 10, pages 301–340. Wiley- Interscience, Hoboken, NJ, 2nd edition, 2006. doi: 10.1002/047174882X.ch10

  30. [38]

    McAfee, Laya Sleiman, Leon Derczynski, Luis Vega, Maer Rodrigues de Melo, Makesh Nar- simhan Sreedhar, Marcin Chochowski, Mark Cai, Markus Kliegl, Marta M

    Nvidia Aarti Basant, Abhijit Khairnar, Ab- hijit Paithankar, Abhinav Khattar, Adi Ren- duchintala, Adi Renduchintala, Aditya Malte, Akhiad Bercovich, Akshay Hazare, Alejandra Rico, Aleksander Ficek, Alex Kondratenko, Alex Shaposhnikov, Ali Taghibakhshi, Amelia Barton, Ameya Ma...

  31. [39]

    Towards vqa models that can read

    Amanpreet Singh, Vivek Natarajan, Meet Shah, YuJiang, XinleiChen, DhruvBatra, DeviParikh, and Marcus Rohrbach. Towards vqa models that can read. InThe IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019

  32. [40]

    Minesh Mathew, Dimosthenis Karatzas, and C.V. Jawahar. Docvqa: Adatasetforvqaondocument images, 2020. URL https://arxiv.org/abs/ 2007.00398

  33. [41]

    Minesh Mathew, Viraj Bagal, Rubén Tito, Di- mosthenis Karatzas, Ernest Valveny, and C.V. Jawahar. Infographicvqa. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV), pages 1697–1706, January 2022

  34. [42]

    OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024

    Yuliang Liu, Zhang Li, Mingxin Huang, Biao Yang, Wenwen Yu, Chunyuan Li, Xu-Cheng Yin, Cheng-Lin Liu, Lianwen Jin, and Xi- ang Bai. OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024. ISSN 1869-1919...

  35. [43]

    Ocrbench v2: An improved bench- mark for evaluating large multimodal models on visual text localization and reasoning, 2025

    Ling Fu, Zhebin Kuang, Jiajun Song, Mingxin Huang, Biao Yang, Yuzhe Li, Linghao Zhu, Qidi Luo, Xinyu Wang, Hao Lu, Zhang Li, Guozhi Tang, Bin Shan, Chunhui Lin, Qi Liu, Binghong Wu, Hao Feng, Hao Liu, Can Huang, Jingqun Tang, Wei Chen, Lianwen Jin, Yuliang Liu, and Xiang Bai. ...

  36. [44]

    A diagram is worth a dozen images

    Aniruddha Kembhavi, Mike Salvato, Eric Kolve, Minjoon Seo, Hannaneh Hajishirzi, and Ali Farhadi. A diagram is worth a dozen images. InEuropean Conference on Computer Vision (ECCV), pages 235–251. Springer, 2016

  37. [45]

    ChartQA: A benchmark for question answering about charts with visual and logical reasoning

    Ahmed Masry, Xuan Long Do, Jia Qing Tan, Shafiq Joty, and Enamul Hoque. ChartQA: A benchmark for question answering about charts with visual and logical reasoning. InFindings of the Association for Computational Linguis- tics: ACL 2022, pages 2263–2279, Dublin, Ire- land, May ...

  38. [46]

    findings-acl.177

    URL https://aclanthology.org/2022. findings-acl.177

  39. [47]

    Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

    Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, Cong Wei, BotaoYu, RuibinYuan, RenliangSun, Ming Yin, Boyuan Zheng, Zhenzhu Yang, Yibo Liu, Wenhao Huang, Huan Sun, Yu Su, and Wenhu Chen. Mmmu: A mass...

  40. [48]

    Seed-bench: Benchmarking multimodal large language models

    Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang, Rui Wang, Ruimao Zhang, and Ying Shan. Seed-bench: Benchmarking multimodal large language models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR),pages13299–13308, June 2024

  41. [49]

    Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024

    Haoning Wu, Dongxu Li, Bei Chen, and Junnan Li. Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024. URL https://arxiv.org/abs/2407. 15754

  42. [50]

    Token merging: Your ViT but faster

    Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feichtenhofer, and Judy Hoffman. Token merging: Your ViT but faster. InThe Eleventh International Conference on Learning Representations, 2023. URL https: //openreview.net/forum?id=aZ8qbRkUql

  43. [51]

    Beyond attention or similarity: Maximizing conditional diversity for token prun- ing in mllms.arXiv preprint arXiv:2506.10967, 2025

    Qizhe Zhang, Mengzhen Liu, Lichen Li, Ming Lu, Yuan Zhang, Junwen Pan, Qi She, and Shang- hang Zhang. Beyond attention or similarity: Maximizing conditional diversity for token prun- ing in mllms.arXiv preprint arXiv:2506.10967, 2025. 14 RADIO1D: Elastic Representations for Co...

  44. [52]

    Generalized intersection over union

    Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese. Generalized intersection over union. June 2019

  45. [53]

    Harold W. Kuhn. The hungarian method for the assignment problem.Naval Re- search Logistics Quarterly, 2:83–97, 1955. doi: 10.1002/nav.3800020109. URL https://onlinelibrary.wiley.com/doi/ abs/10.1002/nav.3800020109

  46. [54]

    Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

    Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving.arXiv preprint arXiv:1903.11027, 2019

  47. [55]

    Vision transform- ers need registers

    Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transform- ers need registers. InThe Twelfth Interna- tional Conference on Learning Representations,

  48. [56]

    URL https://openreview.net/forum? id=2dnO3LLiJ1

  49. [57]

    An image is worth 32 tokens for reconstruction and generation

    Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URLhttps://openreview.net/ forum?id=tOXoQPRzPL

  50. [58]

    Net2net: Accelerating learning via knowl- edge transfer, 2016

    Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. Net2net: Accelerating learning via knowl- edge transfer, 2016. URLhttps://arxiv.org/ abs/1511.05641. Presented at ICLR 2016

  51. [59]

    Claude E. Shannon. Coding theorems for a dis- crete source with a fidelity criterion.IRE Na- tional Convention Record, 7(4):142–163, 1959

  52. [60]

    512min_T2

    Zhe Chen, Jiannan Wu, Wenhai Wang, and et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2023. URL https://arxiv.org/abs/ 2312.14238. 15 RADIO1D: Elastic Representations for Condensed Vision Modeling A. RADIO1D Architecture ...

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.