Pith. sign in

REVIEW 4 cited by

Whitening Sentence Representations for Better Semantics and Faster Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.15316 v1 pith:GHP3J7IE submitted 2021-03-29 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords sentencemodelrepresentationrepresentationswhiteningachieveachievedbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training models such as BERT have achieved great success in many natural language processing tasks. However, how to obtain better sentence representation through these pre-training models is still worthy to exploit. Previous work has shown that the anisotropy problem is an critical bottleneck for BERT-based sentence representation which hinders the model to fully utilize the underlying semantic features. Therefore, some attempts of boosting the isotropy of sentence distribution, such as flow-based model, have been applied to sentence representations and achieved some improvement. In this paper, we find that the whitening operation in traditional machine learning can similarly enhance the isotropy of sentence representations and achieve competitive results. Furthermore, the whitening technique is also capable of reducing the dimensionality of the sentence representation. Our experimental results show that it can not only achieve promising performance but also significantly reduce the storage cost and accelerate the model retrieval speed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recovering Latent Structures after Variational Bayesian Variable Selection: Fit Assessment and Factor-Number Selection in Partially Exploratory Factor Analysis

    stat.ME 2026-07 accept novelty 6.0 of 10

    A scale-free gain rule applied to variational ELBO paths recovers true factor dimensionality in partially exploratory factor analysis where raw information criteria over-factor.

  2. Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Prompt-based text embeddings can be truncated to a small fraction of their dimensions with little performance loss on classification and clustering, but retrieval and STS degrade faster; the difference tracks lower in...

  3. S2Sent: Nested Selectivity Aware Sentence Representation Learning

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A lightweight cross-layer fusion module, S2Sent, improves unsupervised sentence embeddings by gating and DCT frequency selection across Transformer blocks.

  4. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    cs.AI 2026-06 unverdicted novelty 2.0 of 10

    A survey-style reference book mapping the full agentic-AI stack from transformer internals to production deployment, with no new research result.

Pith tools