REVIEW 4 cited by
Core Knowledge Deficits in Multi-Modal Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Multi-modal Large Language Models (MLLMs) demonstrate impressive abilities over high-level perception and reasoning, their robustness in the wild remains limited, often falling short on tasks that are intuitive and effortless for humans. We examine the hypothesis that these deficiencies stem from the absence of core knowledge--rudimentary cognitive abilities innate to humans from early childhood. To explore the core knowledge representation in MLLMs, we introduce CoreCognition, a large-scale benchmark encompassing 12 core knowledge concepts grounded in developmental cognitive science. We evaluate 230 models with 11 different prompts, leading to a total of 2,530 data points for analysis. Our experiments uncover four key findings, collectively demonstrating core knowledge deficits in MLLMs: they consistently underperform and show reduced, or even absent, scalability on low-level abilities relative to high-level ones. Finally, we propose Concept Hacking, a novel controlled evaluation method that reveals MLLMs fail to progress toward genuine core knowledge understanding, but instead rely on shortcut learning as they scale.
Forward citations
Cited by 4 Pith papers
-
Vision Language Models Cannot Reason About Physical Transformation
Current VLMs cannot maintain transformation-invariant representations of number, length, volume or size and instead rely on textual invariance priors that reverse on matched non-conserving controls.
-
VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs
A 1,680-question video benchmark shows leading multimodal models lag humans by ~15 points on visual knowledge, and a See-Think-Answer RL-trained model narrows the gap.
-
Large Language Models Show Signs of Alignment with Human Neurocognition During Abstract Reasoning
Only the largest tested LLMs (about 70 billion parameters) match human accuracy on an abstract reasoning task, and the internal geometry of their best layers correlates moderately with human frontal EEG activity.
-
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Dual-temporal VLM guidance injected into a JEPA predictor via multi-layer pyramid features improves hand-manipulation trajectory forecasting over VLM-only and JEPA-only baselines.
Discussion (0). Sign in to comment.