Pith. sign in

REVIEW 5 major objections 6 minor 133 references

D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse

T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A learned policy reuses ViT work across similar video frames, giving up to 2.64x faster embedding generation within a 2% error bound.

desk verdict A real systems result: learned inter-frame ViT reuse plus GPU compaction gives measured 1.8-2.6x speedups within 2% accuracy on three VideoLM tasks, with interpolated baseline throughputs and limited generalization evidence as the main caveats. read the letter →

arxiv 2506.14107 v2 pith:6XS3REX6 submitted 2025-06-17 cs.DC cs.CV

classification cs.DCcs.CV
keywords video-languagemodelsvisiontransformeraccelerationinter-framecomputationreuselearnedgatingGumbel-SoftmaxtrainingGPUstreamcompactionembeddinggenerationvideoqueryengine
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a video-language query engine can generate the visual embeddings consumed by retrieval, question answering, and grounding models much faster by reusing computations across consecutive frames, while keeping error within 2%. The engine, Déjà Vu, combines a learned ViT variant called ReuseViT with system-level compaction so that the reduced FLOPs actually turn into GPU throughput. On three tasks, the paper reports embedding-generation speedups of 1.81x, 2.64x, and 2.54x over full computation, beating the inter-frame-reuse baselines CMC and Eventful Transformer and the image-based DiffRate at matched accuracy. If the claim holds, large-scale video analytics with video-language models becomes substantially cheaper and more practical.

What carries the argument

The key machinery is the pair of lightweight learned modules inside ReuseViT: a two-layer decision MLP that maps per-token cues (cosine similarity to reference frames, class-token attention, reference type, codec metadata) to a binary reuse mask, and a two-layer restoration MLP (hidden size 128) that calibrates the reused QKV/FFN outputs by adding a correction learned from the token difference. Training uses Gumbel-Softmax soft gating to allow gradients through the discrete decisions, a target-reuse-rate loss, and grouped-frame training so the model learns to tolerate error accumulation. On the system side, layer-wise scheduling across frames lets Déjà Vu free cached activations layer by layer (cached memory compaction) and gather active tokens from multiple frames into dense matrices (sparse computation compaction), which is what converts FLOP reductions into measured throughput.

What would settle it

Run ReuseViT on videos with rapid camera motion, frequent scene cuts, or heavy occlusion — content unlike MSR-VTT, How2QA, and NExT-GQA — and compare end-task accuracy against full recomputation: if clips whose patch-level cosine similarity is high still show embedding or task-accuracy error beyond the 2% bound, the input-similarity premise fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that inter-frame computation reuse in ViT-based video-language models can be learned rather than hand-configured, and that the resulting savings can be made real on GPUs. ReuseViT reuses the QKV projection and feed-forward outputs of tokens whose patch-level cosine similarity to the corresponding patch in a past or future reference frame passes a learned gate; a small restoration MLP corrects the residual difference between current and reference tokens. The decision layer consumes cosine similarity, class-token attention, reference-frame type, and codec metadata, and is trained with a Gumbel-Softmax relaxation, a reuse-rate target, and grouped-frame losses that model error accumulation. The paper reports that this configuration reaches the highest accuracy-throughput tradeoff among CMC, Eventful Transformer, and DiffRate, with up to 2.64x embedding-generation speedup within a 2% error bound.

Load-bearing premise

The load-bearing premise is that a token whose patch looks similar to the same patch in a reference frame will also have similar QKV and feed-forward outputs, so reusing those computed values (plus a small learned correction) keeps the final embedding within the promised accuracy bound.

Editorial extensions

If this is right

  • Embedding generation for retrieval, video QA, and video grounding can be accelerated 1.81x, 2.64x, and 2.54x, respectively, while keeping end-task accuracy within 2%.
  • Learnable reuse decisions beat fixed-rate and fixed-threshold reuse policies: Déjà Vu reaches higher accuracy at the same throughput than Eventful Transformer and CMC, and higher throughput than DiffRate at matched accuracy.
  • Layer-wise scheduling with cached-memory compaction and sparse-computation compaction is what turns FLOP savings into GPU speedups; without them, hard gating alone yields only 1.25x on the QA task.
  • Periodic full I-frame recomputation (roughly every 20 frames) bounds long-sequence error accumulation with less than 5% overhead.
  • Only the small decision and restoration modules need training; the pretrained ViT backbone stays frozen, so deployment does not require re-tuning or storing large model weights.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reuse criterion is input-space cosine similarity; a natural stress test is footage with camera motion or scene cuts where patches remain similar but higher-level content changes, since the paper's evaluation covers three curated video corpora.
  • Because attention layers are excluded from reuse, the speedups should shrink as token counts rise (e.g., 336px or 518px ViTs), where attention becomes a larger FLOP share; extending reuse to attention or sparse-attention kernels would be the next lever.
  • The compaction techniques are described independently of ReuseViT's learned gating, so they could plausibly be bolted onto any sparse ViT accelerator, not just the one evaluated here.
  • The 2% error bound is defined per task accuracy, not per embedding; users who need exact embeddings or who serve adversarial queries would need a fallback path that recomputes.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper proposes Déjà Vu, a video-language query engine that accelerates ViT embedding generation by reusing QKV and FFN computations across consecutive frames. It introduces ReuseViT, which learns per-token reuse decisions through a decision layer, calibrates reused values with a restoration layer, and is trained with a Gumbel-Softmax relaxation and a loss combining cosine similarity to the original output with a target reuse rate. System-level contributions include layer-wise scheduling, cached memory compaction, and stream compaction to convert FLOPs savings into GPU throughput. Evaluations on three VideoLM tasks—MSR-VTT retrieval with CLIP4Clip, How2QA QA with FrozenBiLM, and NExT-GQA grounding with TempCLIP—report throughput improvements up to 1.81x, 2.64x, and 2.54x within a claimed 2% error bound, with comparisons to CMC, Eventful Transformer, and DiffRate.

Significance. If the results hold, Déjà Vu addresses a real bottleneck in video-language analytics by co-designing a learned reuse mechanism with GPU-oriented compaction techniques. The experiments are performed on a commodity RTX 3090 with standard datasets, and the paper explicitly separates FLOPs savings from achieved throughput, which is a methodological strength. The ablation study in Section 7.6 cleanly isolates the contributions of gating, sparse compaction, and memory compaction. The artifact is publicly available, and the training keeps the ViT backbone frozen, which simplifies deployment. However, the comparative advantage over state-of-the-art baselines is weakened by interpolated baseline throughputs, and some evaluation metrics are partly circular or not reported in exact numeric form. The central measurements for Déjà Vu itself are credible, but the claims as currently stated need stronger support or more careful scoping.

major comments (5)
  1. [Section 7.3, Figures 10(d)-(f)] The throughput speedups attributed to CMC and Eventful Transformer are not measured: the paper states that 'we interpolate their throughput assuming our compaction techniques applied.' These interpolated numbers are then used in the abstract and Section 1 to conclude that Déjà Vu outperforms the state of the art (1.81x/2.64x/2.54x versus 1.32x/2.08x/2.20x). Because the authors' compaction techniques may interact differently with CMC's and Eventful's reuse patterns, this assumption is load-bearing and not validated. I ask that the baselines be implemented and measured with the same compaction stack, or that the throughput comparison be explicitly labeled as an estimate with the FLOPs/accuracy comparison presented as the primary evidence.
  2. [Abstract, Section 7] The headline 'within 2% error bound' is not substantiated by an explicit accuracy table. The text reports throughput numbers but not the exact end-task accuracies of the unmodified model and of each Déjà Vu configuration at the operating points used in Figures 10(a)-(f). Since the operating point is chosen through R_target in Eq. 15, the reader needs the actual accuracy drops (e.g., R@5 for MSR-VTT, multiple-choice accuracy for How2QA, GQA@Accuracy for NExT-GQA) to verify the 2% claim. Please add a table reporting these values together with the corresponding R_target for each reported point.
  3. [Section 7.7, Figure 14] The embedding-quality axis in Figure 14 is cosine similarity to the original output, which is exactly the objective optimized through L_sim in Eq. 13. Consequently, the ablation conclusions drawn from Figure 14 partly reflect the model's success in optimizing its training loss rather than an independent measure of quality. To make the design-choice comparisons load-bearing, I request that the same configurations be evaluated with end-task accuracy or another held-out metric, or that the figure be repositioned as reporting satisfaction of the training objective.
  4. [Section 3.3 (Eq. 1), Section 4.2 (Eq. 13), Section 8] The reuse criterion rests on the premise that input-space cosine similarity and the difference ΔR_i predict whether QKV/FFN outputs can be safely reused or restored. The paper does not report any per-token correlation analysis between these signals and the actual post-restoration output error, and the evaluation is limited to three in-distribution benchmarks. Because the 2% error bound is an operating point selected through R_target, not an architectural guarantee, I would like to see a per-token correlation plot or a deliberate domain-shift experiment (e.g., fast camera motion or occlusion) to support transferability. If such evidence is unavailable, the abstract and introduction should explicitly scope the claim to the three evaluated tasks.
  5. [Sections 4.2, 4.3, 6.3] Several hyperparameters that determine the reported tradeoff are not reported: α and R_target in Eq. 15, the Gumbel-Softmax temperature schedule, the I-frame reset period (Section 6.3 mentions 'every twentieth frame' as an example), and the six-frame grouping pattern from Section 4.3. Without these values, it is difficult to reproduce the operating points or to assess how much the 2% error bound depends on hyperparameter tuning. Please include a reproducibility table with the exact values used in the evaluation.
minor comments (6)
  1. [Section 4.1, Eq. (11)] The notation 'GumbelSoftmax(MLP_decision(v))' is ambiguous for binary decisions; if a two-class softmax is intended, the paper should specify how the two logits map to M_soft, or describe the binary concrete distribution formulation.
  2. [Sections 4.3 and 6.3] The training grouping uses six frames in the pattern 1-5-9-13-11-12, whereas online inference collects 'four consecutive frames' per segment; please clarify the relationship between the training grouping and the inference grouping, and whether the periodic I-frame reset every 20 frames is consistent with the trained segment structure.
  3. [Section 6.2] The sentence 'During training, Only the two lightweight modules are trained' has a capitalization error, and the paper does not specify the GPU or wall-clock time used for training beyond the statement that convergence typically occurs within an hour.
  4. [Section 3.3, Eq. (4)] The notation 'M_i ∈ 0,1' should be 'M_i ∈ {0,1}', and the sign convention for the decision-layer output d_i should be stated more explicitly.
  5. [Section 7.1] For the DiffRate baseline, the paper states 'we adapted the policy for VLP models' but does not describe the adaptation; please provide details or a pointer to the adapted implementation in the artifact.
  6. [Section 7.3] The statement that 'only configurations that yield an actual improvement are shown' could hide unfavorable operating points; please specify how many configurations were evaluated and how many are omitted.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: headline speedups are measured on external end-task benchmarks; the internal cosine-similarity ablation is not load-bearing.

full rationale

The central derivation is empirical and self-contained against external benchmarks. ReuseViT's decision/restoration layers are trained with the similarity loss L_sim (Eq. 13) and reuse loss (Eq. 14), and the claimed 1.81x/2.64x/2.54x speedups "within a 2% error bound" are measured by running CLIP4Clip, FrozenBiLM, and TempCLIP on MSR-VTT, How2QA, and NExT-GQA and comparing task accuracy to the unmodified models; no fitted parameter is renamed as a prediction. The user-set R_target in Eq. 15 merely selects an operating point, and the error bound is measured, not assumed. The reliance on Eq. 1's cosine similarity as a reuse-safety signal is a modeling assumption validated only on the three evaluated distributions; this is an empirical-evidence limitation, and Section 8 explicitly leaves broader task generalization open, so it is not a circular derivation. The only internal metric that coincides with a training objective is Figure 14's "cosine similarity" axis, which is equivalent to 1 - L_sim from Eq. 13; that makes the ablation a training-quality diagnostic rather than independent evidence, but it is not load-bearing for the headline claims. Self-citations (CoVA [36], LVS [49]) are background/related-work only and are not load-bearing.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on standard ML machinery (Gumbel-Softmax reparameterization, stream compaction, frozen CLIP/ViT backbones), on domain assumptions about video redundancy validated only empirically on three datasets, and on one explicit interpolation assumption for the baseline comparison. The headline operating point is selected via the user-set target reuse rate rather than derived. No new physical or metaphysical entities are introduced; the only new objects are trained network modules with in-paper ablation evidence.

free parameters (5)
  • R_target (target reuse rate, Eq. 15) = Task-dependent; e.g., 61% in the video-QA ablation; up to ~96% in Figure 10
    User-set knob in the loss that determines the operating point; the headline 'within 2% error' speedups are obtained by choosing R_target, so the 2% bound is an enforced target rather than an emergent property.
  • Alpha (loss weight in Eq. 15) = Not reported
    Balances similarity loss and reuse penalty; no value, search range, or sensitivity analysis is given in the paper.
  • Frame grouping pattern (six frames: 1-5-9-13-11-12) = Six frames, three segments
    Chosen empirically ('we found that 6 frame groupings gave us the best empirical tradeoff'); affects error-accumulation training and training stability.
  • I-frame reset period = Every 20 frames (under 5% overhead)
    Periodic full recomputation breaks error propagation; the period is chosen to fit a 5% overhead budget, so it is tuned rather than derived.
  • Gumbel-Softmax temperature schedule = Not specified
    Annealing from soft to hard gating is described qualitatively (Section 4.1); the exact schedule is absent, so an independent reimplementation may behave differently.
assumptions (5)
  • domain assumption Cosine similarity of input tokens between frames is a valid signal for reusability of QKV/FFN outputs (Eq. 1).
    The gating mechanism assumes tokens similar in input space can have their expensive projections reused with only a small MLP correction (Eqs. 8-9). Validated only empirically on three datasets; no guarantee beyond them.
  • domain assumption Inter-frame redundancy at 2 FPS sampling is sufficient for large reuse without accuracy loss.
    All evaluations sample frames at 2 FPS (Section 7.1); the claimed gains depend on consecutive frames being similar at that rate, which motivates frame reordering but is not guaranteed for arbitrary video.
  • ad hoc to paper Déjà Vu's compaction techniques would give CMC and Eventful Transformer the interpolated throughputs in Figure 10(d)-(f).
    Section 7.3 states 'we interpolate their throughput assuming our compaction techniques applied'; if the techniques do not transfer, the comparative speedup claims are optimistic.
  • standard math Gumbel-Softmax with temperature annealing approximates hard gating well enough for the trained policy to transfer at inference.
    Standard reparameterization (Eq. 11); the annealing schedule is described only qualitatively (Section 4.1), so transfer relies on an unreported schedule choice.
  • domain assumption Attention layers remain a small fraction of FLOPs at the evaluated resolutions (257 tokens per frame for CLIP ViT-B/16 and ViT-L/14).
    Only QKV and FFN layers are reused. Section 8 states attention reaches up to 23.5% of FLOPs at higher token counts (DINOv2-ViT-G/14 at 518px), which is outside the evaluated regime.

how reviews work

0 comments
Cite this review

Pith. "Pith review of D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse." pith.science (2026). https://pith.science/paper/6XS3REX6

@misc{pith2026250614107,
  author       = {Pith},
  title        = {Pith review of: D\'ej\`a Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6XS3REX6}},
  note         = {Machine review of arXiv:2506.14107}
}
read the original abstract

Recently, Video-Language Models (VideoLMs) have demonstrated remarkable capabilities, offering significant potential for flexible and powerful video query systems. These models typically rely on Vision Transformers (ViTs), which process video frames individually to extract visual embeddings. However, generating embeddings for large-scale videos requires ViT inferencing across numerous frames, posing a major hurdle to real-world deployment and necessitating solutions for integration into scalable video data management systems. This paper introduces D\'ej\`a Vu, a video-language query engine that accelerates ViT-based VideoLMs by reusing computations across consecutive frames. At its core is ReuseViT, a modified ViT model specifically designed for VideoLM tasks, which learns to detect inter-frame reuse opportunities, striking an effective balance between accuracy and reuse. Although ReuseViT significantly reduces computation, these savings do not directly translate into performance gains on GPUs. To overcome this, D\'ej\`a Vu integrates memory-compute joint compaction techniques that convert the FLOP savings into tangible performance gains. Evaluations on three VideoLM tasks show that D\'ej\`a Vu accelerates embedding generation by up to a 2.64x within a 2% error bound, dramatically enhancing the practicality of VideoLMs for large-scale video analytics.

Figures

Figures reproduced from arXiv: 2506.14107 by the authors.

Figure 1
Figure 1. Overview of video language query systems sup [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FLOPs breakdown across three VideoLM tasks: video [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 5
Figure 5. FLOPs breakdown of core computations within a [PITH_FULL_IMAGE:figures/full_fig_p004_5.png] view at source ↗
Figures from the paper (7 more)
Figure 6
Figure 6. Figure 6: Overview of ReuseViT model operations. 3.3 Criteria for Computations Reuse An essential aspect of ReuseViT is determining when to reuse com￾putations for specific tokens. We introduce two key components: (1) a Decision Layer that decides whether to reuse computations, …
Figure 7
Figure 7. Figure 7: Illustration of our frame-grouping strategy dur [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 9
Figure 9. Figure 9: Sparse Computation Compaction. so sparse computations can be converted into dense forms better suited to GPU acceleration. Layer-wise stream compaction. We implement GPU kernels for stream compaction across different frames within the same segment, as shown in [PITH_F…
Figure 10
Figure 10. Figure 10: Tradeoff space of baselines and Déjà Vu. Y-axis for FLOPs reduction and throughput are normalized to original model with no reuse. use TempCLIP [104] on NExT-GQA [104], with videos averag￾ing 45 seconds. We report GQA@Accuracy, measuring correct answers with correct t…
Figure 11
Figure 11. Figure 11: Breakdown of FLOPs on video retrieval. (a) FLOPs [PITH_FULL_IMAGE:figures/full_fig_p010_11.png]
Figure 13
Figure 13. Figure 13: Ablation study of Déjà Vu’s inference optimization techniques on the video QA task. 7.5 Memory Overhead [PITH_FULL_IMAGE:figures/full_fig_p010_13.png]
Figure 15
Figure 15. Figure 15: Comparison of Déjà Vu’s adaptive reuse strategy versus Eventful Transformer’s static strategy on a video seg￾ment from How2QA over time. with segment lengths up to three segments, while extending to four segments yields slightly inferior results. Thus, a three-segment…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

133 extracted references · 72 canonical work pages

  1. [1]

    Neil Agarwal and Ravi Netravali. 2023. Boggart: Towards General-Purpose acceleration of retrospective video analytics. InNSDI

  2. [2]

    Michael R Anderson, Michael Cafarella, German Ros, and Thomas F Wenisch

  3. [3]

    Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. Vivit: A video vision transformer. InICCV. 6836–6846

  4. [4]

    Jaeho Bang, Gaurav Tarlok Kakkar, Pramod Chunduri, Subrata Mitra, and Joy Arulraj. 2023. Seiden: Revisiting query processing in video database systems. VLDB16, 9 (2023)

  5. [5]

    Favyen Bastani, Songtao He, Arjun Balasingam, Karthik Gopalakrishnan, Mo- hammad Alizadeh, Hari Balakrishnan, Michael Cafarella, Tim Kraska, and Sam Madden. 2020. Miris: Fast object track queries in video. InSIGMOD. 1907–1921

  6. [6]

    Favyen Bastani and Samuel Madden. 2022. OTIF: Efficient tracker pre-processing over large video datasets. Insigmod. 2091–2104

  7. [7]

    Markus Billeter, Ola Olsson, and Ulf Assarsson. 2009. Efficient stream com- paction on wide SIMD many-core architectures. InHPG. 159–166

  8. [8]

    Daniel Bolya, Cheng-Yang Fu, Xiaoliang Dai, Peizhao Zhang, Christoph Feicht- enhofer, and Judy Hoffman. 2023. Token merging: Your vit but faster.ICLR (2023)

Show all 133 references
  1. [9]

    Mark Buckler, Philip Bedoukian, Suren Jayasuriya, and Adrian Sampson. 2018. EVA2: Exploiting Temporal Redundancy in Live Computer Vision. InISCA. 533–546

  2. [10]

    Adrian Bulat, Juan Manuel Perez Rua, Swathikiran Sudhakaran, Brais Martinez, and Georgios Tzimiropoulos. 2021. Space-time mixing attention for video transformer.NeurIPS34 (2021), 19594–19607

  3. [11]

    Jiashen Cao, Karan Sarkar, Ramyad Hadidi, Joy Arulraj, and Hyesoon Kim. 2022. Figo: Fine-grained query optimization in video analytics. InSIGMOD. 559–572

  4. [12]

    Qingqing Cao, Bhargavi Paranjape, and Hannaneh Hajishirzi. 2023. PuMer: Pruning and Merging Tokens for Efficient Vision Language Models. InProceed- ings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 12890–12903

  5. [13]

    Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. 2021. Emerging properties in self-supervised vision transformers. InICCV. 9650–9660

  6. [14]

    Mengzhao Chen, Wenqi Shao, Peng Xu, Mingbao Lin, Kaipeng Zhang, Fei Chao, Rongrong Ji, Yu Qiao, and Ping Luo. 2023. Diffrate: Differentiable compression rate for efficient vision transformers. InProceedings of the IEEE/CVF international conference on computer vision. 17164–17174

  7. [15]

    Yinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen, Lu Yuan, and Zicheng Liu. 2020. Dynamic convolution: Attention over convolution kernels. InCVPR. 11030–11039

  8. [16]

    Feng Cheng, Xizi Wang, Jie Lei, David Crandall, Mohit Bansal, and Gedas Bertasius. 2023. VindLU: A Recipe for Effective Video-and-Language Pretraining. InCVPR

  9. [17]

    Joonmyung Choi, Sanghyeok Lee, Jaewon Chu, Minhyuk Choi, and Hyunwoo J Kim. 2024. vid-TLDR: Training Free Token merging for Light-weight Video Transformer. InCVPR. 18771–18781

  10. [18]

    Xiangxiang Chu, Zhi Tian, Yuqing Wang, Bo Zhang, Haibing Ren, Xiaolin Wei, Huaxia Xia, and Chunhua Shen. 2021. Twins: Revisiting the design of spatial attention in vision transformers.NIPS34 (2021), 9355–9366

  11. [19]

    Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Zhe Wang, and Doo Soon Kim. 2020. TutorialVQA: Question Answering Dataset for Tutorial Videos. InProceedings of the Twelfth Language Resources and Evaluation Conference. 5450–5455

  12. [20]

    Xiangxiang Dai, Peng Yang, Xinyu Zhang, Zhewei Dai, and Li Yu. 2022. RESPIRE: Reducing Spatial–Temporal Redundancy for Efficient Edge-Based Industrial Video Analytics.IEEE Transactions on Industrial Informatics18, 12 (2022), 9324– 9334

  13. [21]

    Maureen Daum, Brandon Haynes, Dong He, Amrita Mazumdar, and Magdalena Balazinska. 2021. TASM: A Tile-Based Storage Manager for Video Analytics. In ICDE. 1775–1786

  14. [22]

    Maureen Daum, Enhao Zhang, Dong He, Stephen Mussmann, Brandon Haynes, Ranjay Krishna, and Magdalena Balazinska. 2023. Vocalexplore: Pay-as-you-go video data exploration and model building.VLDB16, 13 (2023), 4188–4201

  15. [23]

    Shuchisnigdha Deb, Christopher R Hudson, Daniel W Carruth, and Darren Frey

  16. [24]

    Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Kenneth Liu, Zhenglun Kong, Xin Meng, Zhengang Li, Xue Lin, Zhenman Fang, et al. 2023. Heatvit: Hardware-efficient adaptive token pruning for vision transformers. InHPCA. IEEE

  17. [25]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...

  18. [26]

    Matthew Dutson, Yin Li, and Mohit Gupta. 2023. Eventful transformers: lever- aging temporal redundancy in vision transformers. InICCV. 16911–16923

  19. [27]

    Mohsen Fayyaz, Soroush Abbasi Koohpayegani, Farnoush Rezaei Jafari, Sunando Sengupta, Hamid Reza Vaezi Joze, Eric Sommerlade, Hamed Pirsi- avash, and Jürgen Gall. 2022. Adaptive token sampling for efficient vision transformers. InECCV. Springer

  20. [28]

    Wensheng Gan, Zhenlian Qi, Jiayang Wu, and Jerry Chun-Wei Lin. 2023. Large language models in education: Vision and opportunities. InBigData

  21. [29]

    Ujwalla Gawande, Kamal Hajari, and Yogesh Golhar. 2020. Pedestrian Detection and Tracking in Video Surveillance Vystem: Issues, Comprehensive Review, and Challenges.Recent Trends in Computational Intelligence(2020), 1–24

  22. [30]

    Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze. 2021. Levit: a vision transformer in convnet’s clothing for faster inference. InICCV

  23. [31]

    Deepak Gupta and Dina Demner-Fushman. 2022. Overview of the MedVidQA 2022 shared task on medical video question-answering. InProceedings of the 21st Workshop on Biomedical Language Processing

  24. [32]

    Ramyad Hadidi, Jiashen Cao, Matthew Woodward, Michael S Ryoo, and Hye- soon Kim. 2018. Distributed Perception by Collaborative Robots.IEEE Robotics and Automation Letters3, 4 (2018), 3709–3716

  25. [33]

    Brandon Haynes, Maureen Daum, Dong He, Amrita Mazumdar, Magdalena Balazinska, Alvin Cheung, and Luis Ceze. 2021. VSS: A Storage System for Video Analytics. InSIGMOD

  26. [34]

    Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh. 2021. Rethinking spatial dimensions of vision transformers. InICCV

  27. [35]

    Gibbons, and Onur Mutlu

    Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, and Onur Mutlu. 2018. Focus: Querying Large Video Datasets with Low Latency and Low Cost. In OSDI

  28. [36]

    Jinwoo Hwang, Minsu Kim, Daeun Kim, Seungho Nam, Yoonsung Kim, Dohee Kim, Hardik Sharma, and Jongse Park. 2022. CoVA: Exploiting Compressed- Domain Analysis to Accelerate Video Analytics. InATC

  29. [37]

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig. 2021. Scaling up visual and vision-language representation learning with noisy text supervision. InICML

  30. [38]

    Junchen Jiang, Ganesh Ananthanarayanan, Peter Bodík, Siddhartha Sen, and Ion Stoica. 2018. Chameleon: Scalable Adaptation of Video Analytics. InSIGCOMM

  31. [39]

    Chen Ju, Tengda Han, Kunhao Zheng, Ya Zhang, and Weidi Xie. 2022. Prompting visual-language models for efficient video understanding. InECCV. Springer

  32. [40]

    Gaurav Tarlok Kakkar, Jiashen Cao, Aubhro Sengupta, Joy Arulraj, and Hyesoon Kim. 2024. Hydro: Adaptive Query Processing of ML Queries.arXiv preprint arXiv:2403.14902(2024)

  33. [41]

    Daniel Kang, Peter Bailis, and Matei Zaharia. PVLDB. BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video Analytics. In2019

  34. [42]

    Daniel Kang, John Emmons, Firas Abuzaid, Peter Bailis, and Matei Zaharia

  35. [43]

    Daniel Kang, John Guibas, Peter Bailis, Tatsunori Hashimoto, and Matei Zaharia

  36. [44]

    Daniel Kang, Francisco Romero, Peter D Bailis, Christos Kozyrakis, and Matei Zaharia. 2022. VIVA: An End-to-End System for Interactive Video Analytics.. InCIDR

  37. [45]

    Wonjae Kim, Bokyung Son, and Ildoo Kim. 2021. Vilt: Vision-and-language transformer without convolution or region supervision. InICML

  38. [46]

    Chanwut Kittivorawong, Yongming Ge, Yousef Helal, and Alvin Cheung. 2024. Spatialyze: A Geospatial Video Analytics System with Spatial-Aware Optimiza- tions.VLDB(2024)

  39. [47]

    Zhenglun Kong, Peiyan Dong, Xiaolong Ma, Xin Meng, Mengshu Sun, Wei Niu, Xuan Shen, Geng Yuan, Bin Ren, Minghai Qin, et al. 2022. Spvit: Enabling faster vision transformers via soft token pruning.ECCV(2022)

  40. [48]

    Ziliang Lai, Chris Liu, Chenxia Han, Pengfei Zhang, Eric Lo, and Ben Kao. 2022. Everest: A top-k deep video analytics system. InSIGMOD. 2357–2360

  41. [49]

    Yunghee Lee and Jongse Park. 2024. LVS: A Learned Video Storage for Fast and Efficient Video Understanding. InCVPRW

  42. [50]

    Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi

  43. [51]

    Da Li, Zhang Zhang, Kai Yu, Kaiqi Huang, and Tieniu Tan. 2019. ISEE: an intelligent scene exploration and evaluation platform for large-scale visual surveillance.TPDS(2019)

  44. [52]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InICML

  45. [53]

    Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi. 2021. Align before fuse: Vision and language representation learning with momentum distillation.NIPS(2021)

  46. [54]

    Kunchang Li, Yali Wang, Yizhuo Li, Yi Wang, Yinan He, Limin Wang, and Yu Qiao. 2023. Unmasked teacher: Towards training-efficient video foundation models. InICCV

  47. [55]

    Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu

  48. [56]

    Mengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang, Zhou Zhao, Wenqiao Zhang, Jiaxu Miao, Shiliang Pu, and Fei Wu. 2022. Hero: Hierarchical spatio- temporal reasoning with contrastive action correspondence for end-to-end video object grounding. InMM

  49. [57]

    Ruiyuan Li, Zheng Li, Yi Wu, Chao Chen, and Yu Zheng. 2023. Elf: Erasing-based lossless floating-point compression.VLDB16, 7 (2023), 1763–1776

  50. [58]

    Xirui Li, Chao Ma, Xiaokang Yang, and Ming-Hsuan Yang. 2024. VidToMe: Video Token Merging for Zero-Shot Video Editing.CVPR(2024)

  51. [59]

    Yuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang, Guoqing Harry Xu, and Ravi Netravali. 2020. Reducto: On-camera filtering for resource-efficient real-time video analytics. InSIGCOMM

  52. [60]

    Zheng Li, Soroush Ghodrati, Amir Yazdanbakhsh, Hadi Esmaeilzadeh, and Mingu Kang. 2022. Accelerating Attention through Gradient-Based Learned Runtime Pruning. InProceedings of the 49th Annual International Symposium on Computer Architecture(New York, New York)(ISCA ’22). Assoc...

  53. [61]

    Panagiotis Liakos, Katia Papakonstantinopoulou, and Yannis Kotidis. 2022. Chimp: efficient lossless floating point compression for time series databases. VLDB15, 11 (2022), 3058–3070

  54. [62]

    Youwei Liang, Chongjian Ge, Zhan Tong, Yibing Song, Jue Wang, and Pengtao Xie. 2022. Not all patches are what you need: Expediting vision transformers via token reorganizations. InICLR

  55. [63]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. InICCV

  56. [64]

    Huaishao Luo, Lei Ji, Ming Zhong, Yang Chen, Wen Lei, Nan Duan, and Tianrui Li. 2022. CLIP4Clip: An empirical study of CLIP for end to end video clip retrieval and captioning.Neurocomputing(2022)

  57. [65]

    Dmitrii Marin, Jen-Hao Rick Chang, Anurag Ranjan, Anish Prabhu, Moham- mad Rastegari, and Oncel Tuzel. 2021. Token Pooling in Vision Transformers. arXiv:2110.03860 [cs.CV]

  58. [66]

    Oscar Moll, Manuel Favela, Samuel Madden, Vijay Gadepally, and Michael Cafarella. 2023. SeeSaw: interactive ad-hoc search over image databases.PACM- MOD(2023)

  59. [67]

    Jonghwan Mun, Paul Hongsuck Seo, Ilchae Jung, and Bohyung Han. 2017. Marioqa: Answering questions by watching gameplay videos. InICCV

  60. [68]

    Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po- Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabb...

  61. [69]

    Bowen Pan, Rameswar Panda, Camilo Fosco, Chung-Ching Lin, Alex Andonian, Yue Meng, Kate Saenko, Aude Oliva, and Rogerio Feris. 2021. Video Adap- tive Redundancy Reduction. InProceedings of the International Conference on Learning Representations (ICLR)

  62. [70]

    Junting Pan, Adrian Bulat, Fuwen Tan, Xiatian Zhu, Lukasz Dudziak, Hongsheng Li, Georgios Tzimiropoulos, and Brais Martinez. 2022. Edgevits: Competing light-weight cnns on mobile devices with vision transformers. InECCV

  63. [71]

    Mathias Parger, Chengcheng Tang, Thomas Neff, Christopher D Twigg, Cem Keskin, Robert Wang, and Markus Steinberger. 2023. MotionDeltaCNN: Sparse CNN Inference of Frame Differences in Moving Camera Videos with Spherical Buffers and Padded Convolutions. InICCV

  64. [72]

    Mathias Parger, Chengcheng Tang, Christopher D Twigg, Cem Keskin, Robert Wang, and Markus Steinberger. 2022. DeltaCNN: End-to-end CNN inference of sparse frame differences in videos. InCVPR

  65. [73]

    Tuomas Pelkonen, Scott Franklin, Justin Teller, Paul Cavallaro, Qi Huang, Justin Meza, and Kaushik Veeraraghavan. 2015. Gorilla: A fast, scalable, in-memory time series database.VLDB(2015)

  66. [74]

    AJ Piergiovanni, Weicheng Kuo, and Anelia Angelova. 2023. Rethinking video vits: Sparse video tubes for joint image and video learning. InCVPR

  67. [75]

    Alex Poms, Will Crichton, Pat Hanrahan, and Kayvon Fatahalian. 2018. Scanner: Efficient video analysis at scale.TOG(2018)

  68. [76]

    Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023. FACT: FFN-attention Co-optimized transformer architecture with eager correlation prediction. InProceedings of the 50th Annual International Symposium on Compu...

  69. [77]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sand- hini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al

  70. [78]

    Yongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu, Jie Zhou, and Cho-Jui Hsieh. 2021. Dynamicvit: Efficient vision transformers with dynamic token sparsification.NIPS34 (2021)

  71. [79]

    Shuhuai Ren, Sishuo Chen, Shicheng Li, Xu Sun, and Lu Hou. 2023. TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Under- standing. InEMNLP

  72. [80]

    Francisco Romero, Caleb Winston, Johann Hauswald, Matei Zaharia, and Chris- tos Kozyrakis. 2023. Zelda: Video analytics using vision-language models.arXiv preprint arXiv:2305.03785(2023)

  73. [81]

    Matthew Russo, Tatsunori Hashimoto, Daniel Kang, Yi Sun, and Matei Zaharia

  74. [82]

    Edward Sanderson and Bogdan J Matuszewski. 2022. FCN-transformer fea- ture fusion for polyp segmentation. InAnnual conference on medical image understanding and analysis. Springer

  75. [83]

    Sheng Shen, Chunyuan Li, Xiaowei Hu, Yujia Xie, Jianwei Yang, Pengchuan Zhang, Zhe Gan, Lijuan Wang, Lu Yuan, Ce Liu, et al. 2022. K-lite: Learning transferable visual models with external knowledge.NIPS(2022)

  76. [84]

    Learning transferable visual models from natural language supervision. InICML. PMLR, 8748–8763

  77. [85]

    Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Woj- ciech Galuba, Marcus Rohrbach, and Douwe Kiela. 2022. Flava: A foundational language and vision alignment model. InCVPR

  78. [86]

    Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, and Xiaoyao Liang. 2024. CMC: Video Transformer Acceleration via CODEC Assisted Matrix Condensing. InASPLOS

  79. [87]

    Zhuoran Song, Feiyang Wu, Xueyuan Liu, Jing Ke, Naifeng Jing, and Xiaoyao Liang. 2020. Vr-dann: Real-time video recognition via decoder-assisted neural network acceleration. InMICRO. IEEE, 698–710

  80. [88]

    Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu. 2022. Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning.NIPS(2022)

  81. [89]

    Abhijit Suprem, Joy Arulraj, Calton Pu, and Joao Ferreira. 2020. ODIN: auto- mated drift detection and recovery in video analytics.VLDB(2020)

  82. [90]

    2023.𝛿LTA: Decou- pling Camera Sampling from Processing to Avoid Redundant Computations in the Vision Pipeline

    Raúl Taranco, José-María Arnau, and Antonio González. 2023.𝛿LTA: Decou- pling Camera Sampling from Processing to Avoid Redundant Computations in the Vision Pipeline. InMICRO

  83. [91]

    Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. NeurIPS(2022)

  84. [92]

    Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023. Flexgen: High-throughput generative inference of large language models with a single gpu. InICML

  85. [93]

    Toshiaki Wakatsuki, Sekitoshi Kanai, and Yasuhiro Fujiwara. 2021. Accelerate Inference of CNNs for Video Analysis While Preserving Exactness Exploiting Activation Sparsity. InProceedings of Machine Learning and Systems 3 (MLSys)

  86. [94]

    Hanrui Wang, Zhekai Zhang, and Song Han. 2021. Spatten: Efficient sparse attention architecture with cascade token and head pruning. InHPCA

  87. [95]

    Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023. Videomae v2: Scaling video masked autoencoders with dual masking. InCVPR

  88. [96]

    Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2021. Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. InICCV

  89. [97]

    Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2022. Pvt v2: Improved baselines with pyramid vision transformer.Computational Visual Media8, 3 (2022)

  90. [98]

    Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E Gonzalez. 2018. Skipnet: Learning dynamic routing in convolutional networks. InECCV

  91. [99]

    Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. 2024. Internvid: A large-scale video-text dataset for multimodal understanding and generation.ICLR(2024)

  92. [100]

    Andeep S Toor, Harry Wechsler, and Michele Nappi. 2019. Biometric surveillance using visual question answering.Pattern Recognition Letters(2019)

  93. [101]

    Renzhi Wu, Pramod Chunduri, Ali Payani, Xu Chu, Joy Arulraj, and Kexin Rong. 2024. SketchQL: Video Moment Querying with a Visual Query Interface. Proceedings of the ACM on Management of Data(2024)

  94. [102]

    Zuxuan Wu, Tushar Nagarajan, Abhishek Kumar, Steven Rennie, Larry S Davis, Kristen Grauman, and Rogerio Feris. 2018. Blockdrop: Dynamic inference paths in residual networks. InCVPR

  95. [103]

    Zuxuan Wu, Caiming Xiong, Yu-Gang Jiang, and Larry S Davis. 2019. Liteeval: A coarse-to-fine framework for resource efficient video recognition.NeurIPS (2019)

  96. [104]

    Junbin Xiao, Angela Yao, Yicong Li, and Tat-Seng Chua. 2024. Can i trust your answer? visually grounded video question answering. InCVPR

  97. [105]

    Ziyang Xiao, Dongxiang Zhang, Zepeng Li, Sai Wu, Kian-Lee Tan, and Gang Chen. 2023. DoveDB: A Declarative and Low-Latency Video Database.VLDB (2023)

  98. [106]

    Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016. Msr-vtt: A large video description dataset for bridging video and language. InCVPR

  99. [107]

    Tiantu Xu, Luis Materon Botelho, and Felix Xiaozhu Lin. 2019. Vstore: A data store for analytics on large videos. InProceedings of the Fourteenth EuroSys Conference 2019. 1–17

  100. [108]

    Yi Wang, Kunchang Li, Yizhuo Li, Yinan He, Bingkun Huang, Zhiyu Zhao, Hongjie Zhang, Jilan Xu, Yi Liu, Zun Wang, et al. 2024. Internvideo: General video foundation models via generative and discriminative learning.ECCV (2024)

  101. [109]

    Zhuangdi Xu, Gaurav Tarlok Kakkar, Joy Arulraj, and Umakishore Ramachan- dran. 2022. EVA: A symbolic approach to accelerating exploratory video ana- lytics with materialized views. InSIGMOD

  102. [111]

    Antoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev, and Cordelia Schmid

  103. [112]

    Shusheng Yang, Xinggang Wang, Yu Li, Yuxin Fang, Jiemin Fang, Wenyu Liu, Xun Zhao, and Ying Shan. 2022. Temporally efficient vision transformer for video instance segmentation. InCVPR

  104. [113]

    Seungjae Yoo, Hangyeol Kim, and Joo-Young Kim. 2024. AdapTiV: Sign- Similarity Based Image-Adaptive Token Merging for Vision Transformer Accel- eration. InMICRO. IEEE

  105. [114]

    Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021. Florence: A new foundation model for computer vision.arXiv(2021)

  106. [115]

    Mu Yuan, Lan Zhang, Xuanke You, and Xiang-Yang Li. 2023. PacketGame: Multi- Stream Packet Gating for Concurrent Video Inference at Scale. InSIGCOMM

  107. [116]

    Yifan Xu, Zhijie Zhang, Mengdan Zhang, Kekai Sheng, Ke Li, Weiming Dong, Liqing Zhang, Changsheng Xu, and Xing Sun. 2022. Evo-vit: Slow-fast token evolution for dynamic vision transformer. InAAAI

  108. [117]

    Hang Zhang, Xin Li, and Lidong Bing. 2023. Video-LLaMA: An Instruction- tuned Audio-Visual Language Model for Video Understanding. InEMNLP Demo

  109. [118]

    Shengyu Zhang, Ziqi Tan, Jin Yu, Zhou Zhao, Kun Kuang, Jie Liu, Jingren Zhou, Hongxia Yang, and Fei Wu. 2020. Poet: Product-oriented video captioner for e-commerce. InACM MM

  110. [119]

    Just ask: Learning to answer questions from millions of narrated videos. InICCV

  111. [120]

    Zhipeng Zhang, Xinglin Hou, Kai Niu, Zhongzhen Huang, Tiezheng Ge, Yun- ing Jiang, Qi Wu, and Peng Wang. 2022. Attract me to buy: Advertisement copywriting generation with multimodal multi-structured information.arXiv preprint arXiv:2205.03534(2022)

  112. [121]

    InNIPS, S

    Zero-Shot Video Question Answering via Frozen Bidirectional Language Models. InNIPS, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.)

  113. [122]

    Yue Zhao, Ishan Misra, Philipp Krähenbühl, and Rohit Girdhar. 2023. Learning video representations from large language models. InCVPR

  114. [123]

    Tianxiong Zhong, Zhiwei Zhang, Guo Lu, Ye Yuan, Yu-Ping Wang, and Guoren Wang. 2023. TVM: A Tile-based Video Management Framework.VLDB(2023)

  115. [124]

    Andong Zhu, Sheng Zhang, Xiaohang Shi, Ke Cheng, Hesheng Sun, and San- glu Lu. 2024. Crucio: End-to-End Coordinated Spatio-Temporal Redundancy Elimination for Fast Video Analytics. InINFOCOM

  116. [126]

    Freedman

    Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J. Freedman. 2017. Live Video Analytics at Scale with Approximation and Delay-Tolerance. InNSDI

  117. [129]

    Yuhao Zhang and Arun Kumar. 2020. Panorama: A Data System for Unbounded Vocabulary Querying over Video. InVLDB

  118. [131]

    Zhe Zhang, Chunyu Wang, Weichao Qiu, Wenhu Qin, and Wenjun Zeng. 2020. AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild.CoRRabs/2010.13302 (2020). arXiv:2010.13302 https://arxiv.org/abs/ 2010.13302

  119. [2017]

    In PVLDB

    NoScope: Optimizing Neural Network Queries over Video at Scale. In PVLDB

  120. [2018]

    InProceedings of the Human Factors and Ergonomics Society Annual Meeting

    Pedestrians Receptivity in Autonomous Vehicles: Exploring a Video-based Assessment. InProceedings of the Human Factors and Ergonomics Society Annual Meeting

  121. [2019]

    Physical representation-based predicate optimization for a visual analytics database. InICDE. IEEE, 1466–1477

  122. [2020]

    HERO: Hierarchical Encoder for Video+ Language Omni-representation Pre-training. InEMNLP

  123. [2021]

    Task-agnostic Indexes for Deep Learning-based Queries over Unstructured Data. InPVLDB

  124. [2022]

    Align and prompt: Video-and-language pre-training with entity prompts. InCVPR

  125. [2023]

    VLDB(2023)

    Accelerating Aggregation Queries on Unstructured Streams of Data. VLDB(2023)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.