Pith. sign in

REVIEW 8 cited by

Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04333 v4 pith:S3APCICN submitted 2023-12-07 cs.CL

classification cs.CL
keywords layersllamaknowledgemodeltasksabilitiesanalysisassessing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents an in-depth analysis of Large Language Models (LLMs), focusing on LLaMA, a prominent open-source foundational model in natural language processing. Instead of assessing LLaMA through its generative output, we design multiple-choice tasks to probe its intrinsic understanding in high-order tasks such as reasoning and computation. We examine the model horizontally, comparing different sizes, and vertically, assessing different layers. We unveil several key and uncommon findings based on the designed probing tasks: (1) Horizontally, enlarging model sizes almost could not automatically impart additional knowledge or computational prowess. Instead, it can enhance reasoning abilities, especially in math problem solving, and helps reduce hallucinations, but only beyond certain size thresholds; (2) In vertical analysis, the lower layers of LLaMA lack substantial arithmetic and factual knowledge, showcasing logical thinking, multilingual and recognitive abilities, with top layers housing most computational power and real-world knowledge.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A reinforcement-learning method that rewards refusal over hallucination cuts LLM 'futile reasoning' from ~66-79% to 1-7% on Countdown while roughly preserving accuracy.

  2. BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Combining features from intermediate layers of large vision models improves gait recognition accuracy, and the proposed BiggerGait baseline achieves state-of-the-art results on CCPG and cross-domain benchmarks.

  3. LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing

    cs.CL 2025-05 conditional novelty 6.0 of 10

    LLaMA models encode binary sentiment most strongly in middle layers and emotions in early layers, and truncating the model at the best layer with a probe head yields efficient sentiment classifiers.

  4. Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture

    cs.RO 2025-02 conditional novelty 6.0 of 10

    Linear probes on OpenVLA's Llama backbone decode object and action symbolic states with high accuracy, and the decoded states can be streamed into the DIARC cognitive architecture for real-time monitoring.

  5. Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony

    cs.HC 2026-07 conditional novelty 5.0 of 10

    Group-level EEG dynamic neural synchrony (CorrCA) preferentially tracks the rate of change of continuous arousal and shows valence-dependent structure across four datasets.

  6. Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training

    cs.CL 2025-09 conditional novelty 5.0 of 10

    At 180M parameters and 5B tokens, all layer-wise scaling variants beat the paper's 18-layer uniform baseline, yet the 12-layer uniform baseline remains best.

  7. The Compositional Architecture of Regret in Large Language Models

    cs.CL 2025-06 reject novelty 5.0 of 10

    The paper claims that regret in LLMs is encoded by interacting neuron groups detectable in the final hidden layer, using new S-CDI, RDS, and GIC metrics.

  8. Krutrim LLM: Multilingual Foundational Model for over a Billion People

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Krutrim LLM is a 7B parameter multilingual model trained on 2T tokens with the claimed largest Indic corpus, reporting strong Indic benchmarks and English scores near Llama-2.

Pith tools