Pith. sign in

REVIEW 1 cited by

How Good Are Multi-dimensional Learned Indices? An Experimental Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.05536 v1 pith:XGY6K2S3 submitted 2024-05-09 cs.DB

classification cs.DB
keywords indiceslearnedmulti-dimensionalindexdataevaluationexperimentalgood
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Efficient indexing is fundamental for multi-dimensional data management and analytics. An emerging tendency is to directly learn the storage layout of multi-dimensional data by simple machine learning models, yielding the concept of Learned Index. Compared with the conventional indices used for decades (e.g., kd-tree and R-tree variants), learned indices are empirically shown to be both space- and time-efficient on modern architectures. However, there lacks a comprehensive evaluation of existing multi-dimensional learned indices under a unified benchmark, which makes it difficult to decide the suitable index for specific data and queries and further prevents the deployment of learned indices in real application scenarios. In this paper, we present the first in-depth empirical study to answer the question of how good multi-dimensional learned indices are. Six recently published indices are evaluated under a unified experimental configuration including index implementation, datasets, query workloads, and evaluation metrics. We thoroughly investigate the evaluation results and discuss the findings that may provide insights for future learned index design.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tradeoffs in Processing Queries and Supporting Updates over an ML-Enhanced R-tree

    cs.DB 2025-02 conditional novelty 4.0 of 10

    An ML-enhanced R-tree can process high-overlap range queries up to 5.4X faster than a traditional R-tree, with average query recall up to 99%, but only when the learned model is trained on the same query distribution.

Pith tools