Pith. sign in

REVIEW 1 cited by

A Contextual-Bandit Approach to Personalized News Article Recommendation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1003.0146 v2 pith:DWKR26QE submitted 2010-02-28 cs.LG cs.AIcs.IR

classification cs.LGcs.AIcs.IR
keywords algorithmarticlesbanditcontextuallearningnewspersonalizedservices
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Personalized web services strive to adapt their services (advertisements, news articles, etc) to individual users by making use of both content and user information. Despite a few recent advances, this problem remains challenging for at least two reasons. First, web service is featured with dynamically changing pools of content, rendering traditional collaborative filtering methods inapplicable. Second, the scale of most web services of practical interest calls for solutions that are both fast in learning and computation. In this work, we model personalized recommendation of news articles as a contextual bandit problem, a principled approach in which a learning algorithm sequentially selects articles to serve users based on contextual information about the users and articles, while simultaneously adapting its article-selection strategy based on user-click feedback to maximize total user clicks. The contributions of this work are three-fold. First, we propose a new, general contextual bandit algorithm that is computationally efficient and well motivated from learning theory. Second, we argue that any bandit algorithm can be reliably evaluated offline using previously recorded random traffic. Finally, using this offline evaluation method, we successfully applied our new algorithm to a Yahoo! Front Page Today Module dataset containing over 33 million events. Results showed a 12.5% click lift compared to a standard context-free bandit algorithm, and the advantage becomes even greater when data gets more scarce.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Experiments Under Data Sparse Settings: Applications for Educational Platforms

    cs.LG 2025-01 reject novelty 4.0 of 10

    WAPTS reweights Thompson Sampling draws by empirical success rate to favor high-performing treatments, claiming faster convergence to near-optimal alternatives in sparse educational experiments.

Pith tools