Pith. sign in

REVIEW 7 cited by

Overthinking the Truth: Understanding how Language Models Process False Demonstrations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.09476 v3 pith:CCNL22VO submitted 2023-07-18 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords demonstrationsoverthinkingfalsemodelharmfulheadslayersmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern language models can imitate complex patterns through few-shot learning, enabling them to complete challenging tasks without fine-tuning. However, imitation can also lead models to reproduce inaccuracies or harmful content if present in the context. We study harmful imitation through the lens of a model's internal representations, and identify two related phenomena: "overthinking" and "false induction heads". The first phenomenon, overthinking, appears when we decode predictions from intermediate layers, given correct vs. incorrect few-shot demonstrations. At early layers, both demonstrations induce similar model behavior, but the behavior diverges sharply at some "critical layer", after which the accuracy given incorrect demonstrations progressively decreases. The second phenomenon, false induction heads, are a possible mechanistic cause of overthinking: these are heads in late layers that attend to and copy false information from previous demonstrations, and whose ablation reduces overthinking. Beyond scientific understanding, our results suggest that studying intermediate model computations could be a promising avenue for understanding and guarding against harmful model behaviors.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Verbalizable Representations Form a Global Workspace in Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Language models represent their current reasoning in a small, readable set of verbalizable vectors (the J-space) that functions like a global workspace.

  2. Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models

    cs.CL 2026-07 conditional novelty 7.0 of 10

    Aligned LMs transiently commit to wrong mid-layer preferences that late layers rescue; this wrong-dip predicts structural compression flips, is recipe-specific and trainable, and is distinct from interface failure.

  3. Controlling the Risk of Corrupted Contexts for Language Models via Early-Exiting

    cs.AI 2025-10 conditional novelty 6.0 of 10

    An early-exit rule with a zero-shot fallback, calibrated by Learn-then-Test risk control, keeps the average loss from corrupted in-context demonstrations under a preset bound.

  4. Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Truth directions in LLMs are not universal, emerge only in more capable models, and simple linear probes trained on atomic statements generalize to QA and contextual tasks.

  5. Energy-Guided Decoding for Object Hallucination Mitigation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An energy-guided, training-free decoding rule that chooses the layer with minimal energy reduces object hallucination and yes-bias on several benchmarks.

  6. Improve Decoding Factuality by Token-wise Cross Layer Entropy of Large Language Models

    cs.CL 2025-02 conditional novelty 5.0 of 10

    A new decoding-time method, END, uses per-token cross-layer entropy of prediction growth to boost factual tokens, improving truthfulness and informativeness on hallucination benchmarks.

  7. Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge

    cs.CL 2025-02 conditional novelty 5.0 of 10

    MEMAT combines MEMIT weight edits with optimized attention-head corrections, improving cross-lingual success and magnitude metrics over MEMIT in English and Catalan.

Pith tools