REVIEW 5 major objections 6 minor 133 references
D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A learned policy reuses ViT work across similar video frames, giving up to 2.64x faster embedding generation within a 2% error bound.
desk verdict A real systems result: learned inter-frame ViT reuse plus GPU compaction gives measured 1.8-2.6x speedups within 2% accuracy on three VideoLM tasks, with interpolated baseline throughputs and limited generalization evidence as the main caveats. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the pair of lightweight learned modules inside ReuseViT: a two-layer decision MLP that maps per-token cues (cosine similarity to reference frames, class-token attention, reference type, codec metadata) to a binary reuse mask, and a two-layer restoration MLP (hidden size 128) that calibrates the reused QKV/FFN outputs by adding a correction learned from the token difference. Training uses Gumbel-Softmax soft gating to allow gradients through the discrete decisions, a target-reuse-rate loss, and grouped-frame training so the model learns to tolerate error accumulation. On the system side, layer-wise scheduling across frames lets Déjà Vu free cached activations layer by layer (cached memory compaction) and gather active tokens from multiple frames into dense matrices (sparse computation compaction), which is what converts FLOP reductions into measured throughput.
What would settle it
Run ReuseViT on videos with rapid camera motion, frequent scene cuts, or heavy occlusion — content unlike MSR-VTT, How2QA, and NExT-GQA — and compare end-task accuracy against full recomputation: if clips whose patch-level cosine similarity is high still show embedding or task-accuracy error beyond the 2% bound, the input-similarity premise fails.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that inter-frame computation reuse in ViT-based video-language models can be learned rather than hand-configured, and that the resulting savings can be made real on GPUs. ReuseViT reuses the QKV projection and feed-forward outputs of tokens whose patch-level cosine similarity to the corresponding patch in a past or future reference frame passes a learned gate; a small restoration MLP corrects the residual difference between current and reference tokens. The decision layer consumes cosine similarity, class-token attention, reference-frame type, and codec metadata, and is trained with a Gumbel-Softmax relaxation, a reuse-rate target, and grouped-frame losses that model error accumulation. The paper reports that this configuration reaches the highest accuracy-throughput tradeoff among CMC, Eventful Transformer, and DiffRate, with up to 2.64x embedding-generation speedup within a 2% error bound.
Load-bearing premise
The load-bearing premise is that a token whose patch looks similar to the same patch in a reference frame will also have similar QKV and feed-forward outputs, so reusing those computed values (plus a small learned correction) keeps the final embedding within the promised accuracy bound.
Editorial extensions
If this is right
- Embedding generation for retrieval, video QA, and video grounding can be accelerated 1.81x, 2.64x, and 2.54x, respectively, while keeping end-task accuracy within 2%.
- Learnable reuse decisions beat fixed-rate and fixed-threshold reuse policies: Déjà Vu reaches higher accuracy at the same throughput than Eventful Transformer and CMC, and higher throughput than DiffRate at matched accuracy.
- Layer-wise scheduling with cached-memory compaction and sparse-computation compaction is what turns FLOP savings into GPU speedups; without them, hard gating alone yields only 1.25x on the QA task.
- Periodic full I-frame recomputation (roughly every 20 frames) bounds long-sequence error accumulation with less than 5% overhead.
- Only the small decision and restoration modules need training; the pretrained ViT backbone stays frozen, so deployment does not require re-tuning or storing large model weights.
Reading between the lines
- The reuse criterion is input-space cosine similarity; a natural stress test is footage with camera motion or scene cuts where patches remain similar but higher-level content changes, since the paper's evaluation covers three curated video corpora.
- Because attention layers are excluded from reuse, the speedups should shrink as token counts rise (e.g., 336px or 518px ViTs), where attention becomes a larger FLOP share; extending reuse to attention or sparse-attention kernels would be the next lever.
- The compaction techniques are described independently of ReuseViT's learned gating, so they could plausibly be bolted onto any sparse ViT accelerator, not just the one evaluated here.
- The 2% error bound is defined per task accuracy, not per embedding; users who need exact embeddings or who serve adversarial queries would need a fallback path that recomputes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Déjà Vu, a video-language query engine that accelerates ViT embedding generation by reusing QKV and FFN computations across consecutive frames. It introduces ReuseViT, which learns per-token reuse decisions through a decision layer, calibrates reused values with a restoration layer, and is trained with a Gumbel-Softmax relaxation and a loss combining cosine similarity to the original output with a target reuse rate. System-level contributions include layer-wise scheduling, cached memory compaction, and stream compaction to convert FLOPs savings into GPU throughput. Evaluations on three VideoLM tasks—MSR-VTT retrieval with CLIP4Clip, How2QA QA with FrozenBiLM, and NExT-GQA grounding with TempCLIP—report throughput improvements up to 1.81x, 2.64x, and 2.54x within a claimed 2% error bound, with comparisons to CMC, Eventful Transformer, and DiffRate.
Significance. If the results hold, Déjà Vu addresses a real bottleneck in video-language analytics by co-designing a learned reuse mechanism with GPU-oriented compaction techniques. The experiments are performed on a commodity RTX 3090 with standard datasets, and the paper explicitly separates FLOPs savings from achieved throughput, which is a methodological strength. The ablation study in Section 7.6 cleanly isolates the contributions of gating, sparse compaction, and memory compaction. The artifact is publicly available, and the training keeps the ViT backbone frozen, which simplifies deployment. However, the comparative advantage over state-of-the-art baselines is weakened by interpolated baseline throughputs, and some evaluation metrics are partly circular or not reported in exact numeric form. The central measurements for Déjà Vu itself are credible, but the claims as currently stated need stronger support or more careful scoping.
major comments (5)
- [Section 7.3, Figures 10(d)-(f)] The throughput speedups attributed to CMC and Eventful Transformer are not measured: the paper states that 'we interpolate their throughput assuming our compaction techniques applied.' These interpolated numbers are then used in the abstract and Section 1 to conclude that Déjà Vu outperforms the state of the art (1.81x/2.64x/2.54x versus 1.32x/2.08x/2.20x). Because the authors' compaction techniques may interact differently with CMC's and Eventful's reuse patterns, this assumption is load-bearing and not validated. I ask that the baselines be implemented and measured with the same compaction stack, or that the throughput comparison be explicitly labeled as an estimate with the FLOPs/accuracy comparison presented as the primary evidence.
- [Abstract, Section 7] The headline 'within 2% error bound' is not substantiated by an explicit accuracy table. The text reports throughput numbers but not the exact end-task accuracies of the unmodified model and of each Déjà Vu configuration at the operating points used in Figures 10(a)-(f). Since the operating point is chosen through R_target in Eq. 15, the reader needs the actual accuracy drops (e.g., R@5 for MSR-VTT, multiple-choice accuracy for How2QA, GQA@Accuracy for NExT-GQA) to verify the 2% claim. Please add a table reporting these values together with the corresponding R_target for each reported point.
- [Section 7.7, Figure 14] The embedding-quality axis in Figure 14 is cosine similarity to the original output, which is exactly the objective optimized through L_sim in Eq. 13. Consequently, the ablation conclusions drawn from Figure 14 partly reflect the model's success in optimizing its training loss rather than an independent measure of quality. To make the design-choice comparisons load-bearing, I request that the same configurations be evaluated with end-task accuracy or another held-out metric, or that the figure be repositioned as reporting satisfaction of the training objective.
- [Section 3.3 (Eq. 1), Section 4.2 (Eq. 13), Section 8] The reuse criterion rests on the premise that input-space cosine similarity and the difference ΔR_i predict whether QKV/FFN outputs can be safely reused or restored. The paper does not report any per-token correlation analysis between these signals and the actual post-restoration output error, and the evaluation is limited to three in-distribution benchmarks. Because the 2% error bound is an operating point selected through R_target, not an architectural guarantee, I would like to see a per-token correlation plot or a deliberate domain-shift experiment (e.g., fast camera motion or occlusion) to support transferability. If such evidence is unavailable, the abstract and introduction should explicitly scope the claim to the three evaluated tasks.
- [Sections 4.2, 4.3, 6.3] Several hyperparameters that determine the reported tradeoff are not reported: α and R_target in Eq. 15, the Gumbel-Softmax temperature schedule, the I-frame reset period (Section 6.3 mentions 'every twentieth frame' as an example), and the six-frame grouping pattern from Section 4.3. Without these values, it is difficult to reproduce the operating points or to assess how much the 2% error bound depends on hyperparameter tuning. Please include a reproducibility table with the exact values used in the evaluation.
minor comments (6)
- [Section 4.1, Eq. (11)] The notation 'GumbelSoftmax(MLP_decision(v))' is ambiguous for binary decisions; if a two-class softmax is intended, the paper should specify how the two logits map to M_soft, or describe the binary concrete distribution formulation.
- [Sections 4.3 and 6.3] The training grouping uses six frames in the pattern 1-5-9-13-11-12, whereas online inference collects 'four consecutive frames' per segment; please clarify the relationship between the training grouping and the inference grouping, and whether the periodic I-frame reset every 20 frames is consistent with the trained segment structure.
- [Section 6.2] The sentence 'During training, Only the two lightweight modules are trained' has a capitalization error, and the paper does not specify the GPU or wall-clock time used for training beyond the statement that convergence typically occurs within an hour.
- [Section 3.3, Eq. (4)] The notation 'M_i ∈ 0,1' should be 'M_i ∈ {0,1}', and the sign convention for the decision-layer output d_i should be stated more explicitly.
- [Section 7.1] For the DiffRate baseline, the paper states 'we adapted the policy for VLP models' but does not describe the adaptation; please provide details or a pointer to the adapted implementation in the artifact.
- [Section 7.3] The statement that 'only configurations that yield an actual improvement are shown' could hide unfavorable operating points; please specify how many configurations were evaluated and how many are omitted.
Circularity Check
No significant circularity: headline speedups are measured on external end-task benchmarks; the internal cosine-similarity ablation is not load-bearing.
full rationale
The central derivation is empirical and self-contained against external benchmarks. ReuseViT's decision/restoration layers are trained with the similarity loss L_sim (Eq. 13) and reuse loss (Eq. 14), and the claimed 1.81x/2.64x/2.54x speedups "within a 2% error bound" are measured by running CLIP4Clip, FrozenBiLM, and TempCLIP on MSR-VTT, How2QA, and NExT-GQA and comparing task accuracy to the unmodified models; no fitted parameter is renamed as a prediction. The user-set R_target in Eq. 15 merely selects an operating point, and the error bound is measured, not assumed. The reliance on Eq. 1's cosine similarity as a reuse-safety signal is a modeling assumption validated only on the three evaluated distributions; this is an empirical-evidence limitation, and Section 8 explicitly leaves broader task generalization open, so it is not a circular derivation. The only internal metric that coincides with a training objective is Figure 14's "cosine similarity" axis, which is equivalent to 1 - L_sim from Eq. 13; that makes the ablation a training-quality diagnostic rather than independent evidence, but it is not load-bearing for the headline claims. Self-citations (CoVA [36], LVS [49]) are background/related-work only and are not load-bearing.
Assumptions & free parameters
free parameters (5)
- R_target (target reuse rate, Eq. 15) =
Task-dependent; e.g., 61% in the video-QA ablation; up to ~96% in Figure 10
- Alpha (loss weight in Eq. 15) =
Not reported
- Frame grouping pattern (six frames: 1-5-9-13-11-12) =
Six frames, three segments
- I-frame reset period =
Every 20 frames (under 5% overhead)
- Gumbel-Softmax temperature schedule =
Not specified
assumptions (5)
- domain assumption Cosine similarity of input tokens between frames is a valid signal for reusability of QKV/FFN outputs (Eq. 1).
- domain assumption Inter-frame redundancy at 2 FPS sampling is sufficient for large reuse without accuracy loss.
- ad hoc to paper Déjà Vu's compaction techniques would give CMC and Eventful Transformer the interpolated throughputs in Figure 10(d)-(f).
- standard math Gumbel-Softmax with temperature annealing approximates hard gating well enough for the trained policy to transfer at inference.
- domain assumption Attention layers remain a small fraction of FLOPs at the evaluated resolutions (257 tokens per frame for CLIP ViT-B/16 and ViT-L/14).
Cite this review
Pith. "Pith review of D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse." pith.science (2026). https://pith.science/paper/6XS3REX6
@misc{pith2026250614107,
author = {Pith},
title = {Pith review of: D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse},
year = {2026},
howpublished = {\url{https://pith.science/paper/6XS3REX6}},
note = {Machine review of arXiv:2506.14107}
}
read the original abstract
Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces D\'ej\`a Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, D\'ej\`a Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that D\'ej\`a Vu accelerates embedding generation by up to a 2.64x within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Neil Agarwal and Ravi Netravali. 2023. Boggart: Towards General-Purpose acceleration of retrospective video analytics. InNSDI
2023
-
[2]
Michael R Anderson, Michael Cafarella, German Ros, and Thomas F Wenisch
-
[3]
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. Vivit: A video vision transformer. InICCV. 6836–6846
2021
-
[4]
Jaeho Bang, Gaurav Tarlok Kakkar, Pramod Chunduri, Subrata Mitra, and Joy Arulraj. 2023. Seiden: Revisiting query processing in video database systems. VLDB16, 9 (2023)
2023
-
[5]
Favyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan, Mo- hammad Alizadeh, Hari Balakrishnan, Michael Cafarella, Tim Kraska, and Sam Madden. 2020. Miris: Fast object track queries in video. InSIGMOD. 1907–1921
2020
-
[6]
Favyen Bastani and Samuel Madden. 2022. OTIF: Efficient tracker pre-processing over large video datasets. Insigmod. 2091–2104
2022
-
[7]
Markus Billeter, Ola Olsson, and Ulf Assarsson. 2009. Efficient stream com- paction on wide SIMD many-core architectures. InHPG. 159–166
2009
-
[8]
Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feicht- enhofer, and Judy Hoffman. 2023. Token merging: Your vit but faster.ICLR (2023)
2023
Show all 133 references
-
[9]
Mark Buckler, Philip Bedoukian, Suren Jayasuriya, and Adrian Sampson. 2018. EVA2: Exploiting Temporal Redundancy in Live Computer Vision. InISCA. 533–546
2018
-
[10]
Adrian Bulat, Juan Manuel Perez Rua, Swathikiran Sudhakaran, Brais Martinez, and Georgios Tzimiropoulos. 2021. Space-time mixing attention for video transformer.NeurIPS34 (2021), 19594–19607
2021
-
[11]
Jiashen Cao, Karan Sarkar, Ramyad Hadidi, Joy Arulraj, and Hyesoon Kim. 2022. Figo: Fine-grained query optimization in video analytics. InSIGMOD. 559–572
2022
-
[12]
Qingqing Cao, Bhargavi Paranjape, and Hannaneh Hajishirzi. 2023. PuMer: Pruning and Merging Tokens for Efficient Vision Language Models. InProceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 12890–12903
2023
-
[13]
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging properties in self-supervised vision transformers. InICCV. 9650–9660
2021
-
[14]
Mengzhao Chen, Wenqi Shao, Peng Xu, Mingbao Lin, Kaipeng Zhang, Fei Chao, Rongrong Ji, Yu Qiao, and Ping Luo. 2023. Diffrate: Differentiable compression rate for efficient vision transformers. InProceedings of the IEEE/CVF international conference on computer vision. 17164–17174
2023
-
[15]
Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. 2020. Dynamic convolution: Attention over convolution kernels. InCVPR. 11030–11039
2020
-
[16]
Feng Cheng, Xizi Wang, Jie Lei, David Crandall, Mohit Bansal, and Gedas Bertasius. 2023. VindLU: A Recipe for Effective Video-and-Language Pretraining. InCVPR
2023
-
[17]
Joonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi, and Hyunwoo J Kim. 2024. vid-TLDR: Training Free Token merging for Light-weight Video Transformer. InCVPR. 18771–18781
2024
-
[18]
Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021. Twins: Revisiting the design of spatial attention in vision transformers.NIPS34 (2021), 9355–9366
2021
-
[19]
Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Zhe Wang, and Doo Soon Kim. 2020. TutorialVQA: Question Answering Dataset for Tutorial Videos. InProceedings of the Twelfth Language Resources and Evaluation Conference. 5450–5455
2020
-
[20]
Xiangxiang Dai, Peng Yang, Xinyu Zhang, Zhewei Dai, and Li Yu. 2022. RESPIRE: Reducing Spatial–Temporal Redundancy for Efficient Edge-Based Industrial Video Analytics.IEEE Transactions on Industrial Informatics18, 12 (2022), 9324– 9334
2022
-
[21]
Maureen Daum, Brandon Haynes, Dong He, Amrita Mazumdar, and Magdalena Balazinska. 2021. TASM: A Tile-Based Storage Manager for Video Analytics. In ICDE. 1775–1786
2021
-
[22]
Maureen Daum, Enhao Zhang, Dong He, Stephen Mussmann, Brandon Haynes, Ranjay Krishna, and Magdalena Balazinska. 2023. Vocalexplore: Pay-as-you-go video data exploration and model building.VLDB16, 13 (2023), 4188–4201
2023
-
[23]
Shuchisnigdha Deb, Christopher R Hudson, Daniel W Carruth, and Darren Frey
-
[24]
Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Kenneth Liu, Zhenglun Kong, Xin Meng, Zhengang Li, Xue Lin, Zhenman Fang, et al. 2023. Heatvit: Hardware-efficient adaptive token pruning for vision transformers. InHPCA. IEEE
2023
-
[25]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...
2021
-
[26]
Matthew Dutson, Yin Li, and Mohit Gupta. 2023. Eventful transformers: lever- aging temporal redundancy in vision transformers. InICCV. 16911–16923
2023
-
[27]
Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsi- avash, and Jürgen Gall. 2022. Adaptive token sampling for efficient vision transformers. InECCV. Springer
2022
-
[28]
Wensheng Gan, Zhenlian Qi, Jiayang Wu, and Jerry Chun-Wei Lin. 2023. Large language models in education: Vision and opportunities. InBigData
2023
-
[29]
Ujwalla Gawande, Kamal Hajari, and Yogesh Golhar. 2020. Pedestrian Detection and Tracking in Video Surveillance Vystem: Issues, Comprehensive Review, and Challenges.Recent Trends in Computational Intelligence(2020), 1–24
2020
-
[30]
Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. 2021. Levit: a vision transformer in convnet’s clothing for faster inference. InICCV
2021
-
[31]
Deepak Gupta and Dina Demner-Fushman. 2022. Overview of the MedVidQA 2022 shared task on medical video question-answering. InProceedings of the 21st Workshop on Biomedical Language Processing
2022
-
[32]
Ramyad Hadidi, Jiashen Cao, Matthew Woodward, Michael S Ryoo, and Hye- soon Kim. 2018. Distributed Perception by Collaborative Robots.IEEE Robotics and Automation Letters3, 4 (2018), 3709–3716
2018
-
[33]
Brandon Haynes, Maureen Daum, Dong He, Amrita Mazumdar, Magdalena Balazinska, Alvin Cheung, and Luis Ceze. 2021. VSS: A Storage System for Video Analytics. InSIGMOD
2021
-
[34]
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh. 2021. Rethinking spatial dimensions of vision transformers. InICCV
2021
-
[35]
Gibbons, and Onur Mutlu
Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, and Onur Mutlu. 2018. Focus: Querying Large Video Datasets with Low Latency and Low Cost. In OSDI
2018
-
[36]
Jinwoo Hwang, Minsu Kim, Daeun Kim, Seungho Nam, Yoonsung Kim, Dohee Kim, Hardik Sharma, and Jongse Park. 2022. CoVA: Exploiting Compressed- Domain Analysis to Accelerate Video Analytics. InATC
2022
-
[37]
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. InICML
2021
-
[38]
Junchen Jiang, Ganesh Ananthanarayanan, Peter Bodík, Siddhartha Sen, and Ion Stoica. 2018. Chameleon: Scalable Adaptation of Video Analytics. InSIGCOMM
2018
-
[39]
Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. 2022. Prompting visual-language models for efficient video understanding. InECCV. Springer
2022
-
[40]
Gaurav Tarlok Kakkar, Jiashen Cao, Aubhro Sengupta, Joy Arulraj, and Hyesoon Kim. 2024. Hydro: Adaptive Query Processing of ML Queries.arXiv preprint arXiv:2403.14902(2024)
2024 arXiv
-
[41]
Daniel Kang, Peter Bailis, and Matei Zaharia. PVLDB. BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics. In2019
-
[42]
Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia
-
[43]
Daniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto, and Matei Zaharia
-
[44]
Daniel Kang, Francisco Romero, Peter D Bailis, Christos Kozyrakis, and Matei Zaharia. 2022. VIVA: An End-to-End System for Interactive Video Analytics.. InCIDR
2022
-
[45]
Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. InICML
2021
-
[46]
Chanwut Kittivorawong, Yongming Ge, Yousef Helal, and Alvin Cheung. 2024. Spatialyze: A Geospatial Video Analytics System with Spatial-Aware Optimiza- tions.VLDB(2024)
2024
-
[47]
Zhenglun Kong, Peiyan Dong, Xiaolong Ma, Xin Meng, Mengshu Sun, Wei Niu, Xuan Shen, Geng Yuan, Bin Ren, Minghai Qin, et al. 2022. Spvit: Enabling faster vision transformers via soft token pruning.ECCV(2022)
2022
-
[48]
Ziliang Lai, Chris Liu, Chenxia Han, Pengfei Zhang, Eric Lo, and Ben Kao. 2022. Everest: A top-k deep video analytics system. InSIGMOD. 2357–2360
2022
-
[49]
Yunghee Lee and Jongse Park. 2024. LVS: A Learned Video Storage for Fast and Efficient Video Understanding. InCVPRW
2024
-
[50]
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi
-
[51]
Da Li, Zhang Zhang, Kai Yu, Kaiqi Huang, and Tieniu Tan. 2019. ISEE: an intelligent scene exploration and evaluation platform for large-scale visual surveillance.TPDS(2019)
2019
-
[52]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InICML
2022
-
[53]
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation.NIPS(2021)
2021
-
[54]
Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao. 2023. Unmasked teacher: Towards training-efficient video foundation models. InICCV
2023
-
[55]
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu
-
[56]
Mengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang, Zhou Zhao, Wenqiao Zhang, Jiaxu Miao, Shiliang Pu, and Fei Wu. 2022. Hero: Hierarchical spatio- temporal reasoning with contrastive action correspondence for end-to-end video object grounding. InMM
2022
-
[57]
Ruiyuan Li, Zheng Li, Yi Wu, Chao Chen, and Yu Zheng. 2023. Elf: Erasing-based lossless floating-point compression.VLDB16, 7 (2023), 1763–1776
2023
-
[58]
Xirui Li, Chao Ma, Xiaokang Yang, and Ming-Hsuan Yang. 2024. VidToMe: Video Token Merging for Zero-Shot Video Editing.CVPR(2024)
2024
-
[59]
Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang, Guoqing Harry Xu, and Ravi Netravali. 2020. Reducto: On-camera filtering for resource-efficient real-time video analytics. InSIGCOMM
2020
-
[60]
Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, and Mingu Kang. 2022. Accelerating Attention through Gradient-Based Learned Runtime Pruning. InProceedings of the 49th Annual International Symposium on Computer Architecture(New York, New York)(ISCA ’22). Assoc...
2022
-
[61]
Panagiotis Liakos, Katia Papakonstantinopoulou, and Yannis Kotidis. 2022. Chimp: efficient lossless floating point compression for time series databases. VLDB15, 11 (2022), 3058–3070
2022
-
[62]
Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. 2022. Not all patches are what you need: Expediting vision transformers via token reorganizations. InICLR
2022
-
[63]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. InICCV
2021
-
[64]
Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning.Neurocomputing(2022)
2022
-
[65]
Dmitrii Marin, Jen-Hao Rick Chang, Anurag Ranjan, Anish Prabhu, Moham- mad Rastegari, and Oncel Tuzel. 2021. Token Pooling in Vision Transformers. arXiv:2110.03860 [cs.CV]
2021 arXiv
-
[66]
Oscar Moll, Manuel Favela, Samuel Madden, Vijay Gadepally, and Michael Cafarella. 2023. SeeSaw: interactive ad-hoc search over image databases.PACM- MOD(2023)
2023
-
[67]
Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2017. Marioqa: Answering questions by watching gameplay videos. InICCV
2017
-
[68]
Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po- Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabb...
2024
-
[69]
Bowen Pan, Rameswar Panda, Camilo Fosco, Chung-Ching Lin, Alex Andonian, Yue Meng, Kate Saenko, Aude Oliva, and Rogerio Feris. 2021. Video Adap- tive Redundancy Reduction. InProceedings of the International Conference on Learning Representations (ICLR)
2021
-
[70]
Junting Pan, Adrian Bulat, Fuwen Tan, Xiatian Zhu, Lukasz Dudziak, Hongsheng Li, Georgios Tzimiropoulos, and Brais Martinez. 2022. Edgevits: Competing light-weight cnns on mobile devices with vision transformers. InECCV
2022
-
[71]
Mathias Parger, Chengcheng Tang, Thomas Neff, Christopher D Twigg, Cem Keskin, Robert Wang, and Markus Steinberger. 2023. MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded Convolutions. InICCV
2023
-
[72]
Mathias Parger, Chengcheng Tang, Christopher D Twigg, Cem Keskin, Robert Wang, and Markus Steinberger. 2022. DeltaCNN: End-to-end CNN inference of sparse frame differences in videos. InCVPR
2022
-
[73]
Tuomas Pelkonen, Scott Franklin, Justin Teller, Paul Cavallaro, Qi Huang, Justin Meza, and Kaushik Veeraraghavan. 2015. Gorilla: A fast, scalable, in-memory time series database.VLDB(2015)
2015
-
[74]
AJ Piergiovanni, Weicheng Kuo, and Anelia Angelova. 2023. Rethinking video vits: Sparse video tubes for joint image and video learning. InCVPR
2023
-
[75]
Alex Poms, Will Crichton, Pat Hanrahan, and Kayvon Fatahalian. 2018. Scanner: Efficient video analysis at scale.TOG(2018)
2018
-
[76]
Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023. FACT: FFN-attention Co-optimized transformer architecture with eager correlation prediction. InProceedings of the 50th Annual International Symposium on Compu...
2023
-
[77]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al
-
[78]
Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. 2021. Dynamicvit: Efficient vision transformers with dynamic token sparsification.NIPS34 (2021)
2021
-
[79]
Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu Sun, and Lu Hou. 2023. TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Under- standing. InEMNLP
2023
-
[80]
Francisco Romero, Caleb Winston, Johann Hauswald, Matei Zaharia, and Chris- tos Kozyrakis. 2023. Zelda: Video analytics using vision-language models.arXiv preprint arXiv:2305.03785(2023)
2023 arXiv
-
[81]
Matthew Russo, Tatsunori Hashimoto, Daniel Kang, Yi Sun, and Matei Zaharia
-
[82]
Edward Sanderson and Bogdan J Matuszewski. 2022. FCN-transformer fea- ture fusion for polyp segmentation. InAnnual conference on medical image understanding and analysis. Springer
2022
-
[83]
Sheng Shen, Chunyuan Li, Xiaowei Hu, Yujia Xie, Jianwei Yang, Pengchuan Zhang, Zhe Gan, Lijuan Wang, Lu Yuan, Ce Liu, et al. 2022. K-lite: Learning transferable visual models with external knowledge.NIPS(2022)
2022
-
[84]
Learning transferable visual models from natural language supervision. InICML. PMLR, 8748–8763
-
[85]
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Woj- ciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022. Flava: A foundational language and vision alignment model. InCVPR
2022
-
[86]
Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, and Xiaoyao Liang. 2024. CMC: Video Transformer Acceleration via CODEC Assisted Matrix Condensing. InASPLOS
2024
-
[87]
Zhuoran Song, Feiyang Wu, Xueyuan Liu, Jing Ke, Naifeng Jing, and Xiaoyao Liang. 2020. Vr-dann: Real-time video recognition via decoder-assisted neural network acceleration. InMICRO. IEEE, 698–710
2020
-
[88]
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu. 2022. Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning.NIPS(2022)
2022
-
[89]
Abhijit Suprem, Joy Arulraj, Calton Pu, and Joao Ferreira. 2020. ODIN: auto- mated drift detection and recovery in video analytics.VLDB(2020)
2020
-
[90]
2023.𝛿LTA: Decou- pling Camera Sampling from Processing to Avoid Redundant Computations in the Vision Pipeline
Raúl Taranco, José-María Arnau, and Antonio González. 2023.𝛿LTA: Decou- pling Camera Sampling from Processing to Avoid Redundant Computations in the Vision Pipeline. InMICRO
2023
-
[91]
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. NeurIPS(2022)
2022
-
[92]
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single gpu. InICML
2023
-
[93]
Toshiaki Wakatsuki, Sekitoshi Kanai, and Yasuhiro Fujiwara. 2021. Accelerate Inference of CNNs for Video Analysis While Preserving Exactness Exploiting Activation Sparsity. InProceedings of Machine Learning and Systems 3 (MLSys)
2021
-
[94]
Hanrui Wang, Zhekai Zhang, and Song Han. 2021. Spatten: Efficient sparse attention architecture with cascade token and head pruning. InHPCA
2021
-
[95]
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023. Videomae v2: Scaling video masked autoencoders with dual masking. InCVPR
2023
-
[96]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. InICCV
2021
-
[97]
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2022. Pvt v2: Improved baselines with pyramid vision transformer.Computational Visual Media8, 3 (2022)
2022
-
[98]
Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E Gonzalez. 2018. Skipnet: Learning dynamic routing in convolutional networks. InECCV
2018
-
[99]
Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. 2024. Internvid: A large-scale video-text dataset for multimodal understanding and generation.ICLR(2024)
2024
-
[100]
Andeep S Toor, Harry Wechsler, and Michele Nappi. 2019. Biometric surveillance using visual question answering.Pattern Recognition Letters(2019)
2019
-
[101]
Renzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu, Joy Arulraj, and Kexin Rong. 2024. SketchQL: Video Moment Querying with a Visual Query Interface. Proceedings of the ACM on Management of Data(2024)
2024
-
[102]
Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Rennie, Larry S Davis, Kristen Grauman, and Rogerio Feris. 2018. Blockdrop: Dynamic inference paths in residual networks. InCVPR
2018
-
[103]
Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S Davis. 2019. Liteeval: A coarse-to-fine framework for resource efficient video recognition.NeurIPS (2019)
2019
-
[104]
Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. 2024. Can i trust your answer? visually grounded video question answering. InCVPR
2024
-
[105]
Ziyang Xiao, Dongxiang Zhang, Zepeng Li, Sai Wu, Kian-Lee Tan, and Gang Chen. 2023. DoveDB: A Declarative and Low-Latency Video Database.VLDB (2023)
2023
-
[106]
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. InCVPR
2016
-
[107]
Tiantu Xu, Luis Materon Botelho, and Felix Xiaozhu Lin. 2019. Vstore: A data store for analytics on large videos. InProceedings of the Fourteenth EuroSys Conference 2019. 1–17
2019
-
[108]
Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2024. Internvideo: General video foundation models via generative and discriminative learning.ECCV (2024)
2024
-
[109]
Zhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, and Umakishore Ramachan- dran. 2022. EVA: A symbolic approach to accelerating exploratory video ana- lytics with materialized views. InSIGMOD
2022
-
[111]
Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid
-
[112]
Shusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang, Jiemin Fang, Wenyu Liu, Xun Zhao, and Ying Shan. 2022. Temporally efficient vision transformer for video instance segmentation. InCVPR
2022
-
[113]
Seungjae Yoo, Hangyeol Kim, and Joo-Young Kim. 2024. AdapTiV: Sign- Similarity Based Image-Adaptive Token Merging for Vision Transformer Accel- eration. InMICRO. IEEE
2024
-
[114]
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021. Florence: A new foundation model for computer vision.arXiv(2021)
2021
-
[115]
Mu Yuan, Lan Zhang, Xuanke You, and Xiang-Yang Li. 2023. PacketGame: Multi- Stream Packet Gating for Concurrent Video Inference at Scale. InSIGCOMM
2023
-
[116]
Yifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng, Ke Li, Weiming Dong, Liqing Zhang, Changsheng Xu, and Xing Sun. 2022. Evo-vit: Slow-fast token evolution for dynamic vision transformer. InAAAI
2022
-
[117]
Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction- tuned Audio-Visual Language Model for Video Understanding. InEMNLP Demo
2023
-
[118]
Shengyu Zhang, Ziqi Tan, Jin Yu, Zhou Zhao, Kun Kuang, Jie Liu, Jingren Zhou, Hongxia Yang, and Fei Wu. 2020. Poet: Product-oriented video captioner for e-commerce. InACM MM
2020
-
[119]
Just ask: Learning to answer questions from millions of narrated videos. InICCV
-
[120]
Zhipeng Zhang, Xinglin Hou, Kai Niu, Zhongzhen Huang, Tiezheng Ge, Yun- ing Jiang, Qi Wu, and Peng Wang. 2022. Attract me to buy: Advertisement copywriting generation with multimodal multi-structured information.arXiv preprint arXiv:2205.03534(2022)
2022 arXiv
-
[121]
InNIPS, S
Zero-Shot Video Question Answering via Frozen Bidirectional Language Models. InNIPS, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.)
-
[122]
Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar. 2023. Learning video representations from large language models. InCVPR
2023
-
[123]
Tianxiong Zhong, Zhiwei Zhang, Guo Lu, Ye Yuan, Yu-Ping Wang, and Guoren Wang. 2023. TVM: A Tile-based Video Management Framework.VLDB(2023)
2023
-
[124]
Andong Zhu, Sheng Zhang, Xiaohang Shi, Ke Cheng, Hesheng Sun, and San- glu Lu. 2024. Crucio: End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics. InINFOCOM
2024
-
[126]
Freedman
Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J. Freedman. 2017. Live Video Analytics at Scale with Approximation and Delay-Tolerance. InNSDI
2017
-
[129]
Yuhao Zhang and Arun Kumar. 2020. Panorama: A Data System for Unbounded Vocabulary Querying over Video. InVLDB
2020
-
[131]
Zhe Zhang, Chunyu Wang, Weichao Qiu, Wenhu Qin, and Wenjun Zeng. 2020. AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild.CoRRabs/2010.13302 (2020). arXiv:2010.13302 https://arxiv.org/abs/ 2010.13302
2020 arXiv
-
[2017]
In PVLDB
NoScope: Optimizing Neural Network Queries over Video at Scale. In PVLDB
-
[2018]
InProceedings of the Human Factors and Ergonomics Society Annual Meeting
Pedestrians Receptivity in Autonomous Vehicles: Exploring a Video-based Assessment. InProceedings of the Human Factors and Ergonomics Society Annual Meeting
-
[2019]
Physical representation-based predicate optimization for a visual analytics database. InICDE. IEEE, 1466–1477
-
[2020]
HERO: Hierarchical Encoder for Video+ Language Omni-representation Pre-training. InEMNLP
-
[2021]
Task-agnostic Indexes for Deep Learning-based Queries over Unstructured Data. InPVLDB
-
[2022]
Align and prompt: Video-and-language pre-training with entity prompts. InCVPR
-
[2023]
VLDB(2023)
Accelerating Aggregation Queries on Unstructured Streams of Data. VLDB(2023)
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.