REVIEW 8 cited by
Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper presents an in-depth analysis of Large Language Models (LLMs), focusing on LLaMA, a prominent open-source foundational model in natural language processing. Instead of assessing LLaMA through its generative output, we design multiple-choice tasks to probe its intrinsic understanding in high-order tasks such as reasoning and computation. We examine the model horizontally, comparing different sizes, and vertically, assessing different layers. We unveil several key and uncommon findings based on the designed probing tasks: (1) Horizontally, enlarging model sizes almost could not automatically impart additional knowledge or computational prowess. Instead, it can enhance reasoning abilities, especially in math problem solving, and helps reduce hallucinations, but only beyond certain size thresholds; (2) In vertical analysis, the lower layers of LLaMA lack substantial arithmetic and factual knowledge, showcasing logical thinking, multilingual and recognitive abilities, with top layers housing most computational power and real-world knowledge.
Forward citations
Cited by 8 Pith papers
-
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
A reinforcement-learning method that rewards refusal over hallucination cuts LLM 'futile reasoning' from ~66-79% to 1-7% on Countdown while roughly preserving accuracy.
-
BiggerGait: Unlocking Gait Recognition with Layer-wise Representations from Large Vision Models
Combining features from intermediate layers of large vision models improves gait recognition accuracy, and the proposed BiggerGait baseline achieves state-of-the-art results on CCPG and cross-domain benchmarks.
-
LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing
LLaMA models encode binary sentiment most strongly in middle layers and emotions in early layers, and truncating the model at the best layer with a probe head yields efficient sentiment classifiers.
-
Probing a Vision-Language-Action Model for Symbolic States and Integration into a Cognitive Architecture
Linear probes on OpenVLA's Llama backbone decode object and action symbolic states with high accuracy, and the decoded states can be streamed into the DIARC cognitive architecture for real-time monitoring.
-
Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony
Group-level EEG dynamic neural synchrony (CorrCA) preferentially tracks the rate of change of continuous arousal and shows valence-dependent structure across four datasets.
-
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
At 180M parameters and 5B tokens, all layer-wise scaling variants beat the paper's 18-layer uniform baseline, yet the 12-layer uniform baseline remains best.
-
The Compositional Architecture of Regret in Large Language Models
The paper claims that regret in LLMs is encoded by interacting neuron groups detectable in the final hidden layer, using new S-CDI, RDS, and GIC metrics.
-
Krutrim LLM: Multilingual Foundational Model for over a Billion People
Krutrim LLM is a 7B parameter multilingual model trained on 2T tokens with the claimed largest Indic corpus, reporting strong Indic benchmarks and English scores near Llama-2.
Discussion (0). Continue with ORCID to comment.