Pith. sign in

REVIEW 1 major objections 5 minor 97 references

D-NOVA claims RAG vector retrieval can be executed entirely inside 3D NAND flash arrays, using a new threshold-sensing metric that eliminates off-array re-ranking.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 17:41 UTC pith:DCQZFWSN

load-bearing objection A credible new in-array search mechanism with a thorough simulation study, but the headline gains rest on an unvalidated multi-wordline sensing mode. the 1 major comments →

arxiv 2607.17538 v1 pith:DCQZFWSN submitted 2026-07-20 cs.AR cs.CLcs.DBcs.ETcs.IR

D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation

classification cs.AR cs.CLcs.DBcs.ETcs.IR
keywords in-storage processing3D NAND flashvector retrievalretrieval-augmented generationinverted file indexsimilarity searchthreshold sensingcontrastive adapter
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper contends that the retrieval bottleneck of RAG — moving billions of embeddings from storage to a processor — can be eliminated by making the NAND array itself perform the search. It proposes D-NOVA, an architecture that runs all three stages of an IVF retrieval pipeline inside 3D NAND using a digital, dual-bound threshold-sensing metric called DTS. Vectors are stored vertically along NAND strings, query values are applied as wordline voltages, and each string's on/off state scores a group of dimensions. The SSD controller only sorts compact 4-bit scores and metadata, never raw embeddings. If the sensing mechanism works on real silicon, the claimed results — up to 41.7x faster and 71x more energy-efficient than CPU retrieval, with throughput 4.3–12.1x above the prior in-storage accelerator — would directly address the dominant RAG cost.

Core claim

D-NOVA's central claim is that similarity search over dense vectors can be decomposed into a sequence of binary threshold checks that a NAND string performs natively: with multiple wordlines driven at query-dependent voltages, a string conducts only if every selected cell's threshold voltage lies below its applied voltage. The DTS metric alternates this upper-bound check with a complement-domain check, producing a tight score while staying fully digital. The architecture maps each embedding along a vertical string, stores a 4-bit complement copy, accumulates per-group pass/fail results as a 'deficit' in the existing page-buffer latches, and forwards only compact scores and metadata to the co

What carries the argument

Dual-Bound Tight Similarity Sensing (DTS), a digital distance proxy built from two complement upper-bound checks. UBS drives each of m wordlines at q_i + alpha and marks a group pass only if all cells conduct; Comp UBS does the same on 4-bit complement values (r' = 15 - r), replacing the loose lower-bound check that would otherwise cause false positives. The serial NAND string implements these as a single on/off decision, and per-window score deficits (S_max - S) are accumulated in the 4-bit data latches of the page buffer, avoiding wide adders. Multi-wordline activation and stage-dependent m (e.g., 1–2 for centroid/re-ranking, 4–8 for coarse search) provide the parallelism knob.

Load-bearing premise

The load-bearing premise is that real 3D NAND arrays can simultaneously drive multiple wordlines at distinct query-dependent voltages, with the serial string conduction faithfully implementing the per-cell threshold AND; the paper validates this with RC/SPICE simulation and prior patents, not with silicon measurement of this exact mode.

What would settle it

Program a real QLC die with known INT4 threshold levels, apply m=4–8 wordline voltages at q_i+alpha with alpha=2 under retention noise, and measure string-level conduction against the idealized AND; if per-cell sensing accuracy collapses below usable recall or multi-WL activation requires timing margins that erase the latency gain, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Raw embedding vectors and partial scores never leave the NAND array; only compact candidate metadata crosses to the controller, removing the re-ranking memory wall.
  • The IVF pipeline's three stages run in-storage: centroid selection, coarse Top-K2, and fine Top-K1, with controller sorting accounting for under 0.1% of latency in the final stage.
  • DTS sensing replaces iterative multi-level QLC reads (7–15 sense operations) with a fixed small number of binary senses, giving SLC-like read latency at QLC density.
  • Stage-aware m-adaptive sensing lets the coarse second stage run at m=8 with recall degradation confined to a few percent, since first and third stages stay narrow.
  • A retrieval-aware adapter trained on DTS-mined hard negatives closes most of the recall gap to FP32 IVF, with only 0.5–1% query-encoding overhead.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the multi-wordline sensing mode is validated on silicon, DTS-style bound checking could generalize beyond RAG to other in-storage filtering and similarity-join workloads that can tolerate a threshold score.
  • The adapter recipe suggests a general pattern: when moving a continuous metric onto discrete in-memory hardware, a small learned projection can absorb the metric mismatch without changing storage-side logic.
  • A testable scaling question the paper leaves open is how DTS recall behaves for very large nprobe and more heterogeneous embeddings; the current recall gap is measured at fixed K2=1000, K1=100.
  • The reported energy benefits rely on WL toggling dominating energy; real dies with different peripheral costs could shift the balance, so per-die measurements of multi-WL activation energy are the next check.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. D-NOVA proposes an in-storage retrieval accelerator for RAG that executes the IVF pipeline inside 3D NAND. Embedding vectors are stored as INT4 values in QLC threshold-voltage levels along vertical NAND strings; the query is applied as wordline voltages, and strings are sensed in pass/fail mode to implement a new distance metric, Dual-Bound Tight Similarity Sensing (DTS). DTS combines an upper-bound check with a complement-domain upper-bound check, accumulates per-string score deficits in the page-buffer data latches, and transfers only compact scores and metadata to the SSD controller. A lightweight contrastive adapter, trained offline with DTS-mined hard negatives, maps encoder outputs into a DTS-friendly space. The three IVF stages (centroid search, coarse Top-K2, fine Top-K1 re-ranking) are claimed to run entirely within the NAND array. Evaluations on six datasets and two encoders with an in-house simulator report up to 41.7x lower latency and 71x lower energy than a CPU IVF baseline, and 4.3-12.1x higher throughput and 1.1-1.26x lower energy than the REIS in-storage baseline, with noise-aware Spectre simulations and m-sweep sensitivity studies.

Significance. The core idea is consequential: if the DTS sensing mode is physically realizable, D-NOVA would remove the re-ranking data-movement bottleneck that dominates REIS, and it would be a rare example of executing a full IVF retrieval pipeline inside NAND. The work also has genuine strengths: strict digital sensing avoids analog current accumulation; the contrastive-adapter co-design is a legitimate way to bridge a discrete threshold metric and pretrained embeddings; the stage-aware m-adaptive sensing and locality-aware mapping are well-motivated; and the noise simulations and m-sweeps are unusually thorough. However, the significance is conditional on three load-bearing issues: the unvalidated assumption of simultaneous per-WL query-adaptive voltages, an internal inconsistency in the DTS scoring equation and the Smax/4-bit deficit argument, and the underspecified adapter evaluation protocol. The paper does not release its simulator or code, which limits reproducibility of the headline numbers.

major comments (1)
  1. [§4.2, §5.1, Table 3] The adapter evaluation is underspecified. The paper does not state whether the queries used to train the adapter and to mine DTS hard negatives are disjoint from the evaluation queries. If DTS retrieval results and ground-truth labels from the reported benchmarks are used during training, the 'near-FP32-IVF' recall and the adapter gains in Table 3 and Fig. 11 are partly fitted and are not an independent accuracy measurement. Please specify the train/validation/test split for each benchmark, the negative-pool source and size, the InfoNCE temperature value, and report recall on held-out queries. If a disjoint split is already used, state it explicitly; if not, the accuracy claims must be re-evaluated.
minor comments (5)
  1. [§4.2 / Fig. 9] The InfoNCE loss equation and surrounding text are garbled with unicode artifacts, and Fig. 9 appears twice with different text. The equation is reconstructible, but the manuscript must be cleaned before publication.
  2. [§3.3 / Fig. 8] Figure 8 is duplicated with corrupted labels. Please replace the second copy and ensure all callouts match the figure contents.
  3. [§5.2 / Table 1] The simulator configuration lists the QLC threshold window and PTM models, but does not give the 16-level VTH distribution parameters or the exact RC model inputs. Adding the full configuration would help reproducibility.
  4. [§6.7] The statement that prior work [58] validates activation of up to 48 wordlines should be qualified: that work uses a uniform read voltage for bulk bitwise operations, which does not validate per-WL query-dependent voltages. The current sentence overstates the support provided by [58].
  5. [§6.6] The end-to-end latency/energy results do not state whether the host-side query adapter inference is included. The paper says adapter overhead is 0.5-1% of encoding time; please state explicitly whether it is counted in the reported numbers.

Circularity Check

0 steps flagged

No significant circularity: DTS is a newly introduced construction, the adapter is a co-design/training component rather than a disguised prediction, and headline speedups rest on an external feasibility assumption, not on a definitional identity.

full rationale

D-NOVA's central derivation is not circular. The DTS metric is explicitly defined as a new threshold-based score (UBS and Comp UBS) tailored for NAND string sensing, and the paper states: "D-NOVA does not implement cosine similarity or L2 distance inside the NAND array; instead, each IVF stage uses DTS as a threshold-based retrieval score." Thus DTS is not a renamed version of cosine/L2, and no equation in the paper reduces to its own input by construction. The contrastive adapter is trained offline using DTS-mined hard negatives and InfoNCE loss; the reported Recall@100 is measured against ground-truth retrieval quality, not against the DTS score itself. Training a query adapter to improve a metric-specific search is a legitimate co-design, not a fitted parameter being relabeled as a prediction. The multi-WL activation support is cited from external patents and Flash-Cosmos, and the paper acknowledges its own validation is via Spectre/RC simulations plus prior silicon for bulk bitwise operations. Even if distinct per-WL query-dependent voltages are not directly validated in real silicon, that is a feasibility gap, not a circular reduction. The only self-citations (FeNOMS, Proxima) appear in background or routing references and are not load-bearing for the central claim. The paper is self-contained in its evaluation against an FP32-IVF CPU baseline and the REIS in-storage baseline, and the reported speedups follow from the assumed DTS sensing mechanism rather than from a definitional tautology.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The central claims rest on three categories of assumptions: the physical feasibility of query-voltage multi-WL sensing in 3D NAND, the adequacy of INT4 threshold-level storage under noise, and the generalization of an adapter trained on DTS-mined negatives. Free parameters alpha and m are tuned; no new physical entities are introduced.

free parameters (3)
  • DTS tolerance alpha = 2 (default; swept 1–3)
    DTS bound width; alpha=1 collapsed under modeled QLC noise, alpha=2 chosen as robust operating point (§6.4).
  • Stage-wise parallel sensing degrees m=(m1,m2,m3) = (2,4,2), (2,8,2), (4,8,4)
    Chosen from recall-vs-m sweep (§6.2, §6.6); affects throughput/recall trade-off.
  • Adapter InfoNCE temperature and negative-pool size = not reported
    The adapter loss in §4.2 depends on these; without them the recall numbers are not fully reproducible.
axioms (4)
  • domain assumption 3D NAND supports activating multiple wordlines simultaneously with distinct read voltages while preserving binary string conductance logic.
    Core DTS mechanism depends on this; supported only by patents and prior multi-WL bitwise work, not by a DTS-specific silicon measurement (§3.2, §6.7).
  • domain assumption INT4 embedding dimensions can be stored as ordered V_TH levels and queried by applying voltage q_i+alpha per WL without significant inter-cell interference.
    DTS assumes a clean mapping from INT4 to threshold levels; noise is simulated, but no device data for this storage mode (§5.2).
  • ad hoc to paper DTS score (threshold-bound pass count) is a sufficient ranking function for RAG after query-side adaptation.
    The paper does not derive DTS from cosine/L2; it trains an adapter to close the gap, so ranking quality is an empirical assumption (§4.2, §6.1).
  • domain assumption Adapter trained with DTS-mined hard negatives generalizes to unseen queries/datasets.
    No train/test query split is described; only a database-expansion stress test is given (§6.1).

pith-pipeline@v1.3.0-alltime-deepseek · 32102 in / 18573 out tokens · 164511 ms · 2026-08-01T17:41:14.936038+00:00 · methodology

0 comments
read the original abstract

Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. However, its dense vector retrieval introduces significant latency and energy overhead, becoming the primary performance bottleneck. Although recent in-storage accelerators aim to reduce data movement, they still rely on host or embedded processors outside the memory, where nearly 70% of the total retrieval time is spent. As a result, they cannot fully overcome the bandwidth limitations, leading to yet another memory bottleneck. To tackle these limitations, we present D-NOVA, a hardware-software co-designed in-storage retrieval accelerator. D-NOVA executes an inverted file (IVF)-based hierarchical retrieval pipeline by deeply embedding the search functionality directly into the NAND memory array. This is achieved by incorporating a new distance metric, Dual-Bound Tight Similarity Sensing (DTS), which is specifically tailored for searching within the NAND string. In addition, we introduce a lightweight contrastive adapter that maps embedding vectors into a DTS-friendly domain, recovering near-software recall while improving performance and energy efficiency. D-NOVA is up to 41.7x faster and 71x more energy-efficient than a CPU baseline, and achieves 12.13x higher throughput while being up to 1.26x more energy-efficient than state-of-the-art in-storage RAG accelerators, demonstrating the potential of fully in-storage vector search for scalable RAG acceleration.

Figures

Figures reproduced from arXiv: 2607.17538 by Chang Eun Song, Mingu Kang, Sumukh Pinge, Sung Eun Kim, Tajana S. Rosing, Tianqi Zhang.

Figure 1
Figure 1. Figure 1: (a) RAG system overview, and (b) end-to-end latency [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: IVF three stages. (a) Centroid search, (b) coarse [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: (a) Movement-centric comparison of representative [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 5
Figure 5. Figure 5: (a) Overall architecture of D-NOVA SSD connected [PITH_FULL_IMAGE:figures/full_fig_p004_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Overview of (a) UBS, (b) LBS, and (c) Comp UBS. [PITH_FULL_IMAGE:figures/full_fig_p005_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: (a) DTS in NAND and ΔScoreDTS accumulation, (b) second stage accumulated score distributions, (c) correspond￾ing deficit (𝑆max − 𝑆) distributions, and (d) first, and (e) third stage deficit distributions under 𝑚=(1, 1, 1). Final Scoring with both metrics: The final dual-bound score is obtained by aggregating UBS and Comp UBS results across all the 𝑚-cell groups as follows: ScoreDTS = 𝐷∑︁ /𝑚 𝑗=1 (UBSscore 𝑗… view at source ↗
Figure 8
Figure 8. Figure 8: Vector mapping and IVF data flow in D-NOVA for the (a) first, (b) second, and (c) third stages. [PITH_FULL_IMAGE:figures/full_fig_p006_8.png] view at source ↗
Figure 10
Figure 10. Figure 10: (a) Overview of cluster grouping, (b) normal vector [PITH_FULL_IMAGE:figures/full_fig_p008_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Recall@25 and Recall@100 for FP32-IVF, DTS111, and DTS242, each with and without the adapter, across six datasets. [PITH_FULL_IMAGE:figures/full_fig_p010_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: (a)-(d) Recall@100 degradation heatmap across different encoders for the [PITH_FULL_IMAGE:figures/full_fig_p011_12.png] view at source ↗
Figure 14
Figure 14. Figure 14: (a) NAND sensing accuracy in QLC and Recall@100 [PITH_FULL_IMAGE:figures/full_fig_p011_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: ECDF and normalized QPS of sub-block accesses [PITH_FULL_IMAGE:figures/full_fig_p011_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Normalized (a) latency and (b) energy breakdown of D-NOVA compared with CPU, NSP, and REIS baselines. [PITH_FULL_IMAGE:figures/full_fig_p012_16.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

97 extracted references · 3 canonical work pages

  1. [1]

    Advanced Micro Devices, Inc. 2025. AMD uProf: Performance Analysis Tool. https://www.amd.com/en/developer/uprof.html

  2. [2]

    Mohiuddin Ahmed, Raihan Seraj, and Syed Mohammed Shamsul Islam. 2020. The k-means algorithm: A comprehensive survey and performance evaluation. Electronics9, 8 (2020), 1295

  3. [3]

    Toluwalope Ajayi et al. 2019. OpenROAD: Toward a Self-Driving, Open-Source Digital Layout Implementation Tool Chain. InDAC

  4. [4]

    2023.AMD EPYC™9554

    AMD. 2023.AMD EPYC™9554. https://www.amd.com/en/products/processors/ server/epyc/4th-generation-9004-and-8004-series/amd-epyc-9554.html Ac- cessed: 2025-11-16

  5. [5]

    AMD Adaptive & Embedded Computing Group and Samsung. 2023. SmartSSD®Computational Storage Drive: Product Brief. https: //www.xilinx.com/publications/product-briefs/xilinx-smartssd-computational- storage-drive-product-brief.pdf. Accessed Nov. 2025

  6. [6]

    Woorham Bae, Sung-Yong Cho, and Deog-Kyoon Jeong. 2021. A 1.93-pJ/Bit PCI Express Gen4 PHY Transmitter with On-Chip Supply Regulators in 28 nm CMOS. Electronics(2021). https://api.semanticscholar.org/CorpusID:234325987

  7. [7]

    Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas

    Rajeev Balasubramonian, Andrew B. Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories.ACM Transactions on Architecture and Code Optimization14, 2 (2017). doi:10.1145/3092639 Accessed: 2025-11-16

  8. [8]

    Haratsch, Yixin Luo, and Onur Mutlu

    Yu Cai, Saugata Ghose, Erich F. Haratsch, Yixin Luo, and Onur Mutlu. 2017. Error Characterization, Mitigation, and Recovery in Flash-Memory-Based Solid-State Drives.Proc. IEEE105, 9 (2017), 1666–1704. doi:10.1109/JPROC.2017.2713127

  9. [9]

    Jianlyu Chen, Nan Wang, Chaofan Li, Bo Wang, Shitao Xiao, Han Xiao, Hao Liao, Defu Lian, and Zheng Liu. 2025. AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Tah...

  10. [10]

    Kangqi Chen, Rakesh Nadig, Manos Frouzakis, Nika Mansouri Ghiasi, Yu Liang, Haiyu Mao, Jisung Park, Mohammad Sadrosadati, and Onur Mutlu. 2025. REIS: A High-Performance and Energy-Efficient Retrieval System with In-Storage Pro- cessing. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing ...

  11. [11]

    Mingkai Chen, Tianhua Han, Cheng Liu, Shengwen Liang, Kuai Yu, Lei Dai, Ziming Yuan, Ying Wang, Lei Zhang, Huawei Li, and Xiaowei Li. 2025. DRIM- ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs. InProceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’25). Associat...

  12. [12]

    Qi Chen, Bing Zhao, Haidong Wang, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang, and Jingdong Wang. 2021. SPANN: Highly-efficient Billion-scale Approximate Nearest Neighbor Search. In35th Conference on Neural Information Processing Systems (NeurIPS 2021)

  13. [13]

    Sitian Chen, Amelie Chi Zhou, Yucheng Shi, Yusen Li, and Xin Yao. 2024. Me- mANNS: Enhancing Billion-Scale ANNS Efficiency with Practical PIM Hardware. arXiv:2410.23805. https://arxiv.org/abs/2410.23805

  14. [15]

    Hwanheechan Choi, Hyungjun Jo, Sangmin Ahn, Insang Han, and Hyungcheol Shin. 2026. Machine learning-based prediction of the impact of random grain boundary Z-interference on Vt distribution in 3-D NAND flash memory.Journal of Computational Electronics25, 1 (2026), 16

  15. [16]

    Myungjun Chun, Jaeyong Lee, Sanggu Lee, Myungsuk Kim, and Jihong Kim

  16. [17]

    Lawrence T Clark, Vinay Vashishtha, Lucian Shifren, Aditya Gujja, Saurabh Sinha, Brian Cline, Chandarasekaran Ramamurthy, and Greg Yeric. 2016. ASAP7: A 7-nm finFET predictive process design kit.Microelectronics Journal53 (July 2016), 105–115

  17. [18]

    Cohere. 2023. wikipedia-2023-11-embed-multilingual-v3. https://huggingface. co/datasets/Cohere/wikipedia-2023-11-embed-multilingual-v3. Hugging Face Datasets

  18. [19]

    Matthijs Douze, Alexandr Guzhva, Chengqi Deng, Jeff Johnson, Gergely Szilvasy, Pierre-Emmanuel Mazaré, Maria Lomeli, Lucas Hosseini, and Hervé Jégou. 2025. THE FAISS LIBRARY.IEEE Transactions on Big Data(2025), 1–17. doi:10.1109/ TBDATA.2025.3618474

  19. [20]

    2021.PM9A3 NVMe PCIe SSD

    Samsung Electronics. 2021.PM9A3 NVMe PCIe SSD. https://semiconductor. samsung.com/ssd/datacenter-ssd/pm9a3/ Accessed: 2025-11-16

  20. [21]

    Keming Fan, Ashkan Moradifirouzabadi, Xiangjin Wu, Zheyu Li, Flavio Ponzina, Anton Persson, Eric Pop, Tajana Rosing, and Mingu Kang. 2024. SpecPCM: A Low-Power PCM-Based In-Memory Computing Accelerator for Full-Stack Mass Spectrometry Analysis.IEEE Journal on Exploratory Solid-State Computational Devices and Circuits10 (2024), 161–169. doi:10.1109/JXCDC.2...

  21. [22]

    Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, and Jie Wu

  22. [23]

    Hakim Hafidi, Mounir Ghogho, Philippe Ciblat, and Ananthram Swami. 2022. Negative sampling strategies for contrastive self-supervised learning of graph representations.Signal Processing190 (2022), 108310

  23. [24]

    Tsutomu Higuchi, Takuyo Kodama, Koji Kato, Ryo Fukuda, Naoya Tokiwa, Mitsuhiro Abe, Teruo Takagiwa, Yuki Shimizu, Junji Musha, Katsuaki Sakurai, Jumpei Sato, Tetsuaki Utsumi, Kazuhide Yoneya, Yasuhiro Suematsu, Toshifumi Hashimoto, Takeshi Hioka, Kosuke Yanagidaira, Masatsugu Kojima, Junya Mat- suno, Kei Shiraishi, Kensuke Yamamoto, Shintaro Hayashi, Tomo...

  24. [25]

    Charles AR Hoare. 1962. Quicksort.The computer journal5, 1 (1962), 10–16

  25. [26]

    Po-Kai Hsu, Weihong Xu, Tajana Rosing, and Shimeng Yu. 2023. An in-storage processing architecture with 3d nand heterogeneous integration for spectra open modification search. InProceedings of the International Symposium on Memory Systems. 1–7

  26. [27]

    Zhengding Hu, Vibha Murthy, Zaifeng Pan, Wanlu Li, Xiaoyi Fang, Yufei Ding, and Yuke Wang. 2025. HedraRAG: Co-Optimizing Generation and Retrieval for Heterogeneous RAG Workflows. InProceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. 623–638

  27. [28]

    Bongjoon Hyun, Taehun Kim, Dongjae Lee, and Minsoo Rhu. 2024. Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). 263–279. doi:10.1109/HPCA57654.2024.00029

  28. [29]

    2023.LPDDR4X/LPDDR4 SDRAM MT53E768M64D4, MT53E1536M64D8, MT53E768M32D2, MT53E1536M32D4 Data Sheet

    Micron Technology Inc. 2023.LPDDR4X/LPDDR4 SDRAM MT53E768M64D4, MT53E1536M64D8, MT53E768M32D2, MT53E1536M32D4 Data Sheet. https://www. 13 MICRO 2026, October 31–November 04, 2026, Athens, Greece Chang Eun Song, Sumukh Pinge, Tianqi Zhang, Sung Eun Kim, Tajana S Rosing, and Mingu Kang mouser.com/datasheet/2/671/z4bm_embedded_lpddr4x_lpddr4-3193428.pdf Rev....

  29. [30]

    Yeonwoo Jeong, Hyunji Cho, Kyuri Park, Youngjae Kim, and Sungyong Park. 2025. CALL: Context-Aware Low-Latency Retrieval in Disk-Based Vector Databases. arXiv preprint arXiv:2509.18670(2025)

  30. [31]

    Wenqi Jiang, Suvinay Subramanian, Cat Graves, Gustavo Alonso, Amir Yazdan- bakhsh, and Vidushi Dadu. 2025. Rago: Systematic performance optimization for retrieval-augmented generation serving. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 974–989

  31. [32]

    Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Shufan Liu, Xuanzhe Liu, and Xin Jin. 2024. Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Systems(2024)

  32. [33]

    Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-scale similarity search with GPUs.IEEE Transactions on Big Data7, 3 (2019), 535–547

  33. [34]

    Norman P Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, et al. 2021. Ten lessons from three generations shaped google’s tpuv4i: Industrial product. In 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 1–14

  34. [35]

    Ali Khakifirooz, Sriram Balasubrahmanyam, Richard Fastow, Kristopher H Gaewsky, Chang Wan Ha, Rezaul Haque, Owen W Jungroth, Steven Law, Alias- gar S Madraswala, Binh Ngo, et al. 2021. 30.2 a 1tb 4b/cell 144-tier floating-gate 3d-nand flash memory with 40mb/s program throughput and 13.8 gb/mm 2 bit density. In2021 IEEE International Solid-State Circuits C...

  35. [36]

    Hyun-Jin Kim, Jeong-Don Lim, Jang-Woo Lee, Dae-Hoon Na, Joon-Ho Shin, Chae-Hoon Kim, Seung-Woo Yu, Ji-Yeon Shin, Seon-Kyoo Lee, Devraj Rajagopal, et al. 2015. 7.6 1gb/s 2tb nand flash multi-chip package with frequency-boosting interface chip. In2015 IEEE International Solid-State Circuits Conference-(ISSCC) Digest of Technical Papers. IEEE, 1–3

  36. [37]

    Ji-Hoon Kim, Yeo-Reum Park, Jaeyoung Do, Soo-Young Ji, and Joo-Young Kim

  37. [38]

    Kana Kudo, Yuta Aiba, Kazuma Hasegawa, Xu Li, Yuichi Sano, and Tomoya Sanuki. 2025. Energy-Efficient In-Memory Computing using 3D Flash Memory with Sequential Multi-Block Activation and Current Control Cell (CC cell). In 2025 IEEE International Memory Workshop (IMW). 1–4. doi:10.1109/IMW61990. 2025.11026979

  38. [39]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural Questions: A Benchmark for Question Answering Research.Tr...

  39. [40]

    Comput.72, 1 (2022), 278–290

    Accelerating large-scale graph-based nearest neighbor search on a compu- tational storage platform.IEEE Trans. Comput.72, 1 (2022), 278–290

  40. [41]

    Joo Hwan Lee, Hui Zhang, Veronica Lagrange, Praveen Krishnamoorthy, Xi- aodong Zhao, and Yang Seok Ki. 2020. SmartSSD: FPGA Accelerated Near-Storage Data Analytics on SSD.IEEE Computer Architecture Letters19, 2 (2020), 110–113. doi:10.1109/LCA.2020.3009347

  41. [42]

    Kyungmin Lee, Gunwook Yoon, Seung Jae Baik, and Myounggon Kang. 2026. Low-Power Stack-Level Programming Enabled by Optimized Dummy Word Line Voltage in 3-D NAND Flash Memory.IEEE Journal of the Electron Devices Society 14 (2026), 102–106. doi:10.1109/JEDS.2026.3659350

  42. [43]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th symposium on operating systems principles. 611–626

  43. [44]

    Nancy Leong, Sachit Chandra, and Hounien Chen. 2008. Random cache read using a double memory. US Patent 7,423,915

  44. [45]

    Yinan Li, Bailu Ding, Ziyun Wei, Lukas M Maas, Momin Al-Ghosien, Spyros Blanas, Nicolas Bruno, Carlo Curino, Matteo Interlandi, Craig Peeper, et al. 2025. Scaling GPU-Accelerated Databases beyond GPU Memory Size.Proceedings of the VLDB Endowment18, 11 (2025), 4518–4531

  45. [46]

    Peter Wung Lee. 2015. NAND array hierarchical bit-line structures for multiple word-line and all-bit-line simultaneous erase, erase-verify, program, program-verify, and read operations. https://patents.google.com/patent/ WO2015013689A2/en

  46. [48]

    Hosam M Mahmoud, Reza Modarres, and Robert T Smythe. 1995. Analysis of quickselect: An algorithm for order statistics.RAIRO-Theoretical Informatics and Applications29, 4 (1995), 255–276

  47. [49]

    Arm Ltd. 2016. Cortex-R8. https://www.arm.com/products/silicon-ip-cpu/cortex- r/cortex-r8. Accessed: 2025-11-16

  48. [50]

    Conrado Martínez and Salvador Roura. 2001. Optimal sampling strategies in quicksort and quickselect.SIAM J. Comput.31, 3 (2001), 683–705

  49. [51]

    Micron Technology

    Inc. Micron Technology. 2025.DDR4 SDRAM. https://www.micron.com/products/ memory/dram-components/ddr4-sdram Accessed: 2025-11-16

  50. [52]

    Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018. WWW’18 Open Challenge: Finan- cial Opinion Mining and Question Answering. InCompanion Proceedings of the The Web Conference 2018. Lyon, France, 1941–1942. doi:10.1145/3184558.3192301

  51. [53]

    Seock-Hwan Noh, Hoyeon Lee, Junkyum Kim, Junsu Im, Jay H Park, Sungjin Lee, Sam H Noh, Yeseong Kim, and Jaeha Kung. 2025. Flexible In-NAND Cryp- tographic Processing for Secure Flash Storage.arXiv preprint arXiv:2508.03866 (2025)

  52. [54]

    Yoshiaki Ogura et al . 2003. Semiconductor memory device and method for selecting multiple word lines. https://patents.google.com/patent/JP2003222422A Laid-open patent application

  53. [55]

    Daehoon Na, Jang-woo Lee, Seon-Kyoo Lee, Hwasuk Cho, Junha Lee, Manjae Yang, Eunjin Song, Anil Kavala, Tongsung Kim, Dong-Su Jang, et al . 2021. A 1.8-Gb/s/pin 16-Tb NAND flash memory multi-chip package with F-chip for high-performance and high-capacity storage.IEEE Journal of Solid-State Circuits 56, 4 (2021), 1129–1140

  54. [56]

    Nikolaos Papandreou, Haralampos Pozidis, Nikolas Ioannou, Thomas Parnell, Roman Pletka, Milos Stanisavljevic, Radu Stoica, Sasa Tomic, Patrick Breen, Gary Tressler, et al. 2020. Open block characterization and read voltage calibration of 3D QLC NAND flash. In2020 IEEE International Reliability Physics Symposium (IRPS). IEEE, 1–6

  55. [57]

    Krishna Parat and Chuck Dennison. 2015. A floating gate based 3D NAND technology with CMOS under array. In2015 IEEE International Electron Devices Meeting (IEDM). IEEE, 3–3

  56. [58]

    Hiroyuki Ootomo, Akira Naruse, Corey Nolet, Ray Wang, Tamas Feher, and Yong Wang. 2024. Cagra: Highly parallel graph construction and approximate nearest neighbor search for gpus. In2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 4236–4247

  57. [59]

    Jisung Park, Myungsuk Kim, Myoungjun Chun, Lois Orosa, Jihong Kim, and Onur Mutlu. 2021. Reducing solid-state drive read latency by optimizing read-retry. InProceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems(Virtual, USA)(ASPLOS ’21). Association for Computing Machinery, New York, ...

  58. [60]

    Advait Parulekar, Liam Collins, Karthikeyan Shanmugam, Aryan Mokhtari, and Sanjay Shakkottai. 2023. Infonce loss provably learns cluster-preserving rep- resentations. InThe Thirty Sixth Annual Conference on Learning Theory. PMLR, 1914–1961

  59. [61]

    Jisung Park, Roknoddin Azizi, Geraldo F Oliveira, Mohammad Sadrosadati, Rakesh Nadig, David Novo, Juan Gómez-Luna, Myungsuk Kim, and Onur Mutlu

  60. [62]

    In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO)

    Flash-cosmos: In-flash bulk bitwise operations using inherent computation capability of nand flash memory. In2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 937–955

  61. [63]

    Derrick Quinn, Mohammad Nouri, Neel Patel, John Salihu, Alireza Salemi, Sukhan Lee, Hamed Zamani, and Mohammad Alian. 2025. Accelerating retrieval- augmented generation. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume

  62. [64]

    D. J. Rosenkrantz, R. E. Stearns, and P. M. Lewis. 1977. An Analysis of Several Heuristics for the Traveling Salesman Problem.SIAM J. Comput.6, 3 (1977), 563–581. doi:10.1137/0206041

  63. [65]

    Sumukh Pinge, Ashkan Moradifirouzabadi, Keming Fan, Prasanna Venkatesan Ravindran, Tanvir H Pantha, Po-Kai Hsu, Zheyu Li, Weihong Xu, Zihan Xia, Flavio Ponzina, et al . 2025. FeNOMS: Enhancing Open Modification Spectral Library Search with In-Storage Processing on Ferroelectric NAND (FeNAND) Flash. In2025 IEEE/ACM International Conference On Computer Aide...

  64. [66]

    Yubin Qin, Yang Wang, Dazheng Deng, Zhiren Zhao, Xiaolong Yang, Leibo Liu, Shaojun Wei, Yang Hu, and Shouyi Yin. 2023. FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation Prediction. InProceedings of the 50th Annual International Symposium on Computer Architecture(Orlando, FL, USA)(ISCA ’23). Association for Computing Machiner...

  65. [67]

    Edward Suh, and Udit Gupta

    Michael Shen, Muhammad Umar, Kiwan Maeng, G. Edward Suh, and Udit Gupta

  66. [68]

    Joobo Shim, Jaewon Oh, Hongchan Roh, Jaeyoung Do, and Sang-Won Lee. 2025. Turbocharging Vector Databases Using Modern SSDs.Proc. VLDB Endow.18, 11 (July 2025), 4710–4722. doi:10.14778/3749646.3749724

  67. [69]

    Sayed Ahmad Salehi. 2022. In-memory Bulk Bitwise Logic Operation for Multi- level Cell Non-volatile Memories. InProceedings of the 2022 International Sympo- sium on Memory Systems. 1–5

  68. [70]

    Eran Sharon et al. 2014. Simultaneous sensing of multiple word-lines and detec- tion of NAND failures. https://patents.google.com/patent/EP2737487A1/en

  69. [71]

    Chang Eun Song, Priyansh Bhatnagar, Zihan Xia, Nam Sung Kim, Tajana S Rosing, and Mingu Kang. 2025. Hybrid SLC-MLC RRAM Mixed-Signal Processing-in- Memory Architecture for Transformer Acceleration via Gradient Redistribution. InProceedings of the 52nd Annual International Symposium on Computer Architec- ture. 1155–1170

  70. [72]

    InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25)

    Hermes: Algorithm-System Co-design for Efficient Retrieval-Augmented Generation At-Scale. InProceedings of the 52nd Annual International Symposium on Computer Architecture (ISCA ’25). Association for Computing Machinery, New York, NY, USA, 958–973. doi:10.1145/3695053.3731076 14 D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Sim...

  71. [73]

    Chang Eun Song, Ashkan Moradifirouzabadi, Tajana Rosing, and Mingu Kang

  72. [74]

    Wonbo Shim, Hongwu Jiang, Xiaochen Peng, and Shimeng Yu. 2021. Architectural Design of 3D NAND Flash based Compute-in-Memory for Inference Engine. In Proceedings of the International Symposium on Memory Systems(Washington, DC, USA)(MEMSYS ’20). Association for Computing Machinery, New York, NY, USA, 77–85. doi:10.1145/3422575.3422779

  73. [75]

    Tinku Singh, Durgesh Kumar Srivastava, and Alok Aggarwal. 2017. A novel approach for CPU utilization on a multicore paradigm using parallel quicksort. In 2017 3rd International Conference on Computational Intelligence & Communication Technology (CICT). IEEE, 1–6

  74. [76]

    Yoshiki Takai, Mamoru Fukuchi, Reika Kinoshita, Chihiro Matsui, and Ken Takeuchi. 2019. Analysis on heterogeneous SSD configuration with quadruple- level cell (QLC) NAND flash memory. In2019 IEEE 11th International Memory Workshop (IMW). IEEE, 1–4

  75. [77]

    Chang Eun Song, Yidong Li, Amardeep Ramnani, Pulkit Agrawal, Purvi Agrawal, Sung-Joon Jang, Sang-Seol Lee, Tajana Rosing, and Mingu Kang. 2024. 52.5 TOPS/W 1.7 GHz Reconfigurable XGBoost Inference Accelerator Based on Modular-Unit-Tree with Dynamic Data and Compute Gating. In2024 IEEE Custom Integrated Circuits Conference (CICC). IEEE, 1–2

  76. [78]

    James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal

  77. [79]

    Bing Tian, Haikun Liu, Zhuohui Duan, Xiaofei Liao, Hai Jin, and Yu Zhang. 2024. Scalable billion-point approximate nearest neighbor search using SmartSSDs. In Proceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference (Santa Clara, CA, USA)(USENIX ATC’24). USENIX Association, USA, Article 69, 16 pages

  78. [80]

    Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu. 2020. MPNet: Masked and Permuted Pre-training for Language Understanding. InAdvances in Neural Information Processing Systems 33 (NeurIPS 2020). https://proceedings.neurips. cc/paper/2020/file/c3a690be93aa602ee2dc0ccab5b7b67e-Paper.pdf

  79. [81]

    Stillmaker and B

    A. Stillmaker and B. Baas. 2017. Scaling equations for the accurate prediction of CMOS device performance from 180 nm to 7 nm.Integration, the VLSI Jour- nal58 (2017), 74–81. http://vcl.ece.ucdavis.edu/pubs/2017.02.VLSIintegration. TechScale/

  80. [82]

    Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou

Showing first 80 references.