Pith. sign in

REVIEW 1 cited by

On Search Engine Evaluation Metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1302.2318 v1 pith:S6MMHA3N submitted 2013-02-10 cs.IR

classification cs.IR
keywords evaluationmetricsenginesearchframeworkindividualmeta-evaluationmetric
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The search engine evaluation research has quite a lot metrics available to it. Only recently, the question of the significance of individual metrics started being raised, as these metrics' correlations to real-world user experiences or performance have generally not been well-studied. The first part of this thesis provides an overview of previous literature on the evaluation of search engine evaluation metrics themselves, as well as critiques of and comments on individual studies and approaches. The second part introduces a meta-evaluation metric, the Preference Identification Ratio (PIR), that quantifies the capacity of an evaluation metric to capture users' preferences. Also, a framework for simultaneously evaluating many metrics while varying their parameters and evaluation standards is introduced. Both PIR and the meta-evaluation framework are tested in a study which shows some interesting preliminary results; in particular, the unquestioning adherence to metrics or their ad hoc parameters seems to be disadvantageous. Instead, evaluation methods should themselves be rigorously evaluated with regard to goals set for a particular study.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Accelerated learning from recommender systems using multi-armed bandit

    cs.IR 2019-08 conditional novelty 4.0 of 10

    A Vrbo team used daily Thompson sampling to rank four recommendation models by click-through rate, but the A/B validation they report is for a previous campaign's winner, not the current one.

Pith tools