Pith. sign in

REVIEW 3 cited by

Surface Form Competition: Why the Highest Probability Answer Isn't Always Right

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08315 v9 pith:JZH7WRVN submitted 2021-04-16 cs.CL

classification cs.CL
keywords probabilitysurfaceanswerchoicecompetitionformmultiplezero-shot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models have shown promising results in zero-shot settings (Brown et al.,2020; Radford et al., 2019). For example, they can perform multiple choice tasks simply by conditioning on a question and selecting the answer with the highest probability. However, ranking by string probability can be problematic due to surface form competition-wherein different surface forms compete for probability mass, even if they represent the same underlying concept, e.g. "computer" and "PC." Since probability mass is finite, this lowers the probability of the correct answer, due to competition from other strings that are valid answers (but not one of the multiple choice options). We introduce Domain Conditional Pointwise Mutual Information, an alternative scoring function that directly compensates for surface form competition by simply reweighing each option according to a term that is proportional to its a priori likelihood within the context of the specific zero-shot task. It achieves consistent gains in zero-shot performance over both calibrated (Zhao et al., 2021) and uncalibrated scoring functions on all GPT-2 and GPT-3 models over a variety of multiple choice datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Primes as Explanans for Emotion in Large Language Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    NSM semantic primes are more recoverable, more causally effective, and behaviorally more interchangeable with emotions than appraisal dimensions in four instruction-tuned LLMs.

  2. Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models

    cs.HC 2025-06 conditional novelty 6.0 of 10

    Interactive Reasoning, instantiated as Hippo, lets users view and edit an LLM's chain-of-thought as a tree, and a 16-person study reports improved perceived control, sense-making, and assumption awareness.

  3. Why Thinking Hurts: Diagnosing and Rectifying Linguistic Inertia in Large Language Models for Recommendation

    cs.IR 2026-02 conditional novelty 5.0 of 10

    Chain-of-thought reasoning degrades semantic-ID recommendation accuracy through 'linguistic inertia,' and a training-free compression-plus-contrastive decoding fix restores and often improves accuracy.

Pith tools