Pith. sign in

REVIEW 1 cited by

Long-Term Value of Exploration: Measurements, Findings and Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.07764 v2 pith:LNFN7BGF submitted 2023-05-12 cs.IR

classification cs.IR
keywords explorationlong-termalgorithmbanditbenefitscontentcorpusdesigns
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Effective exploration is believed to positively influence the long-term user experience on recommendation platforms. Determining its exact benefits, however, has been challenging. Regular A/B tests on exploration often measure neutral or even negative engagement metrics while failing to capture its long-term benefits. We here introduce new experiment designs to formally quantify the long-term value of exploration by examining its effects on content corpus, and connecting content corpus growth to the long-term user experience from real-world experiments. Once established the values of exploration, we investigate the Neural Linear Bandit algorithm as a general framework to introduce exploration into any deep learning based ranking systems. We conduct live experiments on one of the largest short-form video recommendation platforms that serves billions of users to validate the new experiment designs, quantify the long-term values of exploration, and to verify the effectiveness of the adopted neural linear bandit algorithm for exploration.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Policy Design for Two-sided Platforms with Participation Dynamics

    cs.GT 2025-02 conditional novelty 7.0 of 10

    In a two-sided platform model where viewer and provider populations co-evolve, myopic-greedy recommendation is suboptimal under heterogeneous population effects, and a look-ahead policy improves long-term welfare.

Pith tools