REVIEW 9 cited by
Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The widespread adoption of large language models (LLMs) makes it important to recognize their strengths and limitations. We argue that in order to develop a holistic understanding of these systems we need to consider the problem that they were trained to solve: next-word prediction over Internet text. By recognizing the pressures that this task exerts we can make predictions about the strategies that LLMs will adopt, allowing us to reason about when they will succeed or fail. This approach - which we call the teleological approach - leads us to identify three factors that we hypothesize will influence LLM accuracy: the probability of the task to be performed, the probability of the target output, and the probability of the provided input. We predict that LLMs will achieve higher accuracy when these probabilities are high than when they are low - even in deterministic settings where probability should not matter. To test our predictions, we evaluate two LLMs (GPT-3.5 and GPT-4) on eleven tasks, and we find robust evidence that LLMs are influenced by probability in the ways that we have hypothesized. In many cases, the experiments reveal surprising failure modes. For instance, GPT-4's accuracy at decoding a simple cipher is 51% when the output is a high-probability word sequence but only 13% when it is low-probability. These results show that AI practitioners should be careful about using LLMs in low-probability situations. More broadly, we conclude that we should not evaluate LLMs as if they are humans but should instead treat them as a distinct type of system - one that has been shaped by its own particular set of pressures.
Forward citations
Cited by 9 Pith papers
-
Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives
Combining language-model generation with rule-based selection reproduces several pragmatic phenomena, but the language models only worked reliably as idea generators, not as judges of formal linguistic properties.
-
Modeling Open-World Cognition as On-Demand Synthesis of Probabilistic Models
A hybrid language-model and probabilistic-program architecture predicts human judgments on novel open-world reasoning vignettes better than language-model-only baselines.
-
Reasoning Strategies in Large Language Models: Can They Follow, Prefer, and Optimize?
Prompting LLMs with distinct reasoning strategies and ensembling their outputs improves accuracy on logical deduction tasks, though not as consistently as the paper claims.
-
Transformers Don't In-Context Learn Least Squares Regression
In-context regression transformers do not approximate OLS: they underperform it even in-distribution, fail on out-of-subspace prompts, and their failures correlate with a low-rank spectral signature in the residual stream.
-
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage?
LLMs perform much worse on causal questions built from post-cutoff news articles, suggesting their apparent causal skill is mostly memorization, and a general-knowledge prompt method only partly closes the gap.
-
Thinking beyond the anthropomorphic paradigm benefits LLM research
Anthropomorphic language and assumptions are common and growing in LLM research, and the authors propose a framework for moving beyond them while keeping what is useful.
-
Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
A new maze-navigation benchmark for LLMs reports that reasoning models outperform standard ones, but the link from performance gaps to a lack of persistent self-awareness is an overreach.
-
Benchmarking and Rethinking Knowledge Editing for Large Language Models
Under autoregressive and sequential editing, parameter-based knowledge editing methods perform poorly, while the retrieval-based SCR baseline consistently outperforms them across datasets and models.
-
Towards a Neurosymbolic Reasoning System Grounded in Schematic Representations
Embodied-LM translates natural-language statements into spatial relations (left, right, inside) as Answer Set Programs, then solves them with Clingo/Z3, scoring 91% on LogicalDeduction.
Discussion (0). Continue with ORCID to comment.