Pith. sign in

REVIEW 8 cited by

Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.06588 v1 pith:AJOUU5FA submitted 2022-06-14 cs.IR cs.LG

classification cs.IRcs.LG
keywords datasetsearchqueriesresultsproductimprovingqueryshopping
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Improving the quality of search results can significantly enhance users experience and engagement with search engines. In spite of several recent advancements in the fields of machine learning and data mining, correctly classifying items for a particular user search query has been a long-standing challenge, which still has a large room for improvement. This paper introduces the "Shopping Queries Dataset", a large dataset of difficult Amazon search queries and results, publicly released with the aim of fostering research in improving the quality of search results. The dataset contains around 130 thousand unique queries and 2.6 million manually labeled (query,product) relevance judgements. The dataset is multilingual with queries in English, Japanese, and Spanish. The Shopping Queries Dataset is being used in one of the KDDCup'22 challenges. In this paper, we describe the dataset and present three evaluation tasks along with baseline results: (i) ranking the results list, (ii) classifying product results into relevance categories, and (iii) identifying substitute products for a given query. We anticipate that this data will become the gold standard for future research in the topic of product search.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Verbalized LLM confidence scores are sparse enough that the interpolation method used to compute AUARC reverses method rankings, and a simple logprobs-weighted digit expectation (verbalization logprobs) outperforms va...

  2. MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents

    cs.AI 2026-07 conditional novelty 7.0 of 10

    MMShopBench, built from real multimodal shopping logs, shows even the best agent satisfies the full request in under two-thirds of cases, and fine-tuning on 900 real-log trajectories substantially closes the gap to pr...

  3. APeB: Benchmarking Personalization Ability of Large Language Model Agents

    cs.AI 2026-07 conditional novelty 6.5 of 10

    LLM agents succeed on refined product queries but fail on early underspecified intents mainly because they underuse noisy histories; APeB measures this gap and VQRA partially closes it.

  4. jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation

    cs.IR 2026-07 conditional novelty 6.0 of 10

    jina-reranker-v3.5, a 0.6B listwise reranker with a 3L2G hybrid attention schedule and three-stage self-distillation, scores 63.20 nDCG@10 on BEIR and runs up to 1.56x faster than its predecessor.

  5. KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    KaLM-Reranker-V1 uses encoder–decoder FBNL with Matryoshka pooling to match Qwen3-class reranking quality at substantially lower online cost.

  6. Generative Representational Learning of Foundation Models for Recommendation

    cs.IR 2025-06 conditional novelty 6.0 of 10

    A single recommendation model with task-aware Mixture of Low-rank Experts and convergence-based sample scheduling beats baselines on a new 13-task benchmark.

  7. SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval

    cs.IR 2025-08 conditional novelty 5.0 of 10

    A contrastive fine-tuning recipe that adds HTML structure and element-masking signals improves long structured document retrieval, with reported MRR@10 gains of about four points on BGE-M3.

  8. K-order Ranking Preference Optimization for Large Language Models

    cs.IR 2025-05 conditional novelty 5.0 of 10

    KPO extends the Plackett-Luce preference model used in DPO to top-K partial rankings, with query-adaptive K and curriculum learning, and reports improved LLM ranking accuracy.

Pith tools