Pith. sign in

REVIEW 3 cited by

Problems with Cosine as a Measure of Embedding Similarity for High Frequency Words

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.05092 v1 pith:XKJVAGM2 submitted 2022-05-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords similaritywordscosinefrequencyhighotherembeddingsword
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cosine similarity of contextual embeddings is used in many NLP tasks (e.g., QA, IR, MT) and metrics (e.g., BERTScore). Here, we uncover systematic ways in which word similarities estimated by cosine over BERT embeddings are understated and trace this effect to training data frequency. We find that relative to human judgements, cosine similarity underestimates the similarity of frequent words with other instances of the same word or other words across contexts, even after controlling for polysemy and other factors. We conjecture that this underestimation of similarity for high frequency words is due to differences in the representational geometry of high and low frequency words and provide a formal argument for the two-dimensional case.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

    cs.CL 2026-07 conditional novelty 5.5 of 10

    Relative Probability Association Metric (RPAM) measures LM associations via softmax-normalized continuation probabilities and correlates strongly with human associations and downstream LM behavior across three models.

  2. FlowETL: An Autonomous Example-Driven Pipeline for Data Engineering

    cs.SE 2025-07 conditional novelty 5.0 of 10

    FlowETL uses LLMs and a small target dataset to automatically infer and apply data-cleaning transformations, reporting high data-quality scores across 14 datasets.

  3. Spiritual-LLM : Gita Inspired Mental Health Therapy In the Era of LLMs

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A Gita-based mental-health dialogue dataset helps small LLMs score higher on spirituality-oriented metrics, but the evaluation loop is largely self-referential.

Pith tools