REVIEW 1 cited by
Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
An important assumption that comes with using LLMs on psycholinguistic data has gone unverified. LLM-based predictions are based on subword tokenization, not decomposition of words into morphemes. Does that matter? We carefully test this by comparing surprisal estimates using orthographic, morphological, and BPE tokenization against reading time data. Our results replicate previous findings and provide evidence that in the aggregate, predictions using BPE tokenization do not suffer relative to morphological and orthographic segmentation. However, a finer-grained analysis points to potential issues with relying on BPE-based tokenization, as well as providing promising results involving morphologically-aware surprisal estimates and suggesting a new method for evaluating morphological prediction.
Forward citations
Cited by 1 Pith paper
-
Large Language Models Are Human-Like Internally
Larger LMs appear human-like internally: when surprisal is read from internal layers, their fit to human reading and EEG data matches or exceeds smaller models, overturning final-layer-only conclusions.
Discussion (0). Continue with ORCID to comment.