Pith. sign in

REVIEW 2 cited by

Deep Lake: a Lakehouse for Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.10785 v2 pith:RKTJIRAC submitted 2022-09-22 cs.DC cs.AIcs.CVcs.DB

classification cs.DCcs.AIcs.CVcs.DB
keywords datadeeplakelearningapplicationsdatasetslakehouselakes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Traditional data lakes provide critical data infrastructure for analytical workloads by enabling time travel, running SQL queries, ingesting data with ACID transactions, and visualizing petabyte-scale datasets on cloud storage. They allow organizations to break down data silos, unlock data-driven decision-making, improve operational efficiency, and reduce costs. However, as deep learning usage increases, traditional data lakes are not well-designed for applications such as natural language processing (NLP), audio processing, computer vision, and applications involving non-tabular datasets. This paper presents Deep Lake, an open-source lakehouse for deep learning applications developed at Activeloop. Deep Lake maintains the benefits of a vanilla data lake with one key difference: it stores complex data, such as images, videos, annotations, as well as tabular data, in the form of tensors and rapidly streams the data over the network to (a) Tensor Query Language, (b) in-browser visualization engine, or (c) deep learning frameworks without sacrificing GPU utilization. Datasets stored in Deep Lake can be accessed from PyTorch, TensorFlow, JAX, and integrate with numerous MLOps tools.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LayoutBench: Performance Benchmarking of Cloud Storage Layouts for Multimedia Data

    cs.DC 2026-07 conditional novelty 6.0 of 10

    LayoutBench benchmarks three cloud storage layouts for multimedia retrieval, finding tar-based packing offers the best latency-cost tradeoff for ImageNet-scale data.

  2. AIMS.au: A Dataset for the Analysis of Modern Slavery Countermeasures in Corporate Statements

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Introduces AIMS.au, a 5,731-statement, sentence-level annotated dataset for detecting disclosures mandated by Australia's Modern Slavery Act, with benchmarks showing fine-tuned models outperform zero-shot LLMs.

Pith tools