Pith. sign in

REVIEW 2 cited by

Cross-Validation for Unsupervised Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 0909.3052 v1 pith:D6VH4J5F submitted 2009-09-16 stat.ME math.STstat.TH

classification stat.MEmath.STstat.TH
keywords cross-validationunsupervisedlearningapplychoosingcomponentscontextscriterion
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Cross-validation (CV) is a popular method for model-selection. Unfortunately, it is not immediately obvious how to apply CV to unsupervised or exploratory contexts. This thesis discusses some extensions of cross-validation to unsupervised learning, specifically focusing on the problem of choosing how many principal components to keep. We introduce the latent factor model, define an objective criterion, and show how CV can be used to estimate the intrinsic dimensionality of a data set. Through both simulation and theory, we demonstrate that cross-validation is a valuable tool for unsupervised learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bi-cross validation for estimating spectral clustering hyper parameters

    stat.ML 2019-08 reject novelty 5.0 of 10

    A method using bi-cross validation on the inverted Laplacian matrix to estimate spectral clustering hyperparameters, demonstrated on simulations and LCLS data, but lacking a proof and failing on one synthetic case.

  2. Fusing heterogeneous data sets

    q-bio.GN 2019-08 conditional novelty 4.0 of 10

    The thesis proposes logistic PCA via non-convex singular value thresholding, generalized SCA for binary and quantitative data, and P-ESCA for multiple mixed-type data sets with structured sparsity to separate common a...

Pith tools