REVIEW 8 cited by
Shopping Queries Dataset: A Large-Scale ESCI Benchmark for Improving Product Search
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Improving the quality of search results can significantly enhance users experience and engagement with search engines. In spite of several recent advancements in the fields of machine learning and data mining, correctly classifying items for a particular user search query has been a long-standing challenge, which still has a large room for improvement. This paper introduces the "Shopping Queries Dataset", a large dataset of difficult Amazon search queries and results, publicly released with the aim of fostering research in improving the quality of search results. The dataset contains around 130 thousand unique queries and 2.6 million manually labeled (query,product) relevance judgements. The dataset is multilingual with queries in English, Japanese, and Spanish. The Shopping Queries Dataset is being used in one of the KDDCup'22 challenges. In this paper, we describe the dataset and present three evaluation tasks along with baseline results: (i) ranking the results list, (ii) classifying product results into relevance categories, and (iii) identifying substitute products for a given query. We anticipate that this data will become the gold standard for future research in the topic of product search.
Forward citations
Cited by 8 Pith papers
-
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification
Verbalized LLM confidence scores are sparse enough that the interpolation method used to compute AUARC reverses method rankings, and a simple logprobs-weighted digit expectation (verbalization logprobs) outperforms va...
-
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
MMShopBench, built from real multimodal shopping logs, shows even the best agent satisfies the full request in under two-thirds of cases, and fine-tuning on 900 real-log trajectories substantially closes the gap to pr...
-
APeB: Benchmarking Personalization Ability of Large Language Model Agents
LLM agents succeed on refined product queries but fail on early underspecified intents mainly because they underuse noisy histories; APeB measures this gap and VQRA partially closes it.
-
jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation
jina-reranker-v3.5, a 0.6B listwise reranker with a 3L2G hybrid attention schedule and three-stage self-distillation, scores 63.20 nDCG@10 on BEIR and runs up to 1.56x faster than its predecessor.
-
KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
KaLM-Reranker-V1 uses encoder–decoder FBNL with Matryoshka pooling to match Qwen3-class reranking quality at substantially lower online cost.
-
Generative Representational Learning of Foundation Models for Recommendation
A single recommendation model with task-aware Mixture of Low-rank Experts and convergence-based sample scheduling beats baselines on a new 13-task benchmark.
-
SEAL: Structure and Element Aware Learning to Improve Long Structured Document Retrieval
A contrastive fine-tuning recipe that adds HTML structure and element-masking signals improves long structured document retrieval, with reported MRR@10 gains of about four points on BGE-M3.
-
K-order Ranking Preference Optimization for Large Language Models
KPO extends the Plackett-Luce preference model used in DPO to top-K partial rankings, with query-adaptive K and curriculum learning, and reports improved LLM ranking accuracy.
Discussion (0). Continue with ORCID to comment.