Pith. sign in

REVIEW 1 cited by

Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17774 v1 pith:NAOWHN2J submitted 2023-10-26 cs.CL

classification cs.CL
keywords tokenizationmorphologicaldataestimatesmorphemesorthographicpredictionsresults
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An important assumption that comes with using LLMs on psycholinguistic data has gone unverified. LLM-based predictions are based on subword tokenization, not decomposition of words into morphemes. Does that matter? We carefully test this by comparing surprisal estimates using orthographic, morphological, and BPE tokenization against reading time data. Our results replicate previous findings and provide evidence that in the aggregate, predictions using BPE tokenization do not suffer relative to morphological and orthographic segmentation. However, a finer-grained analysis points to potential issues with relying on BPE-based tokenization, as well as providing promising results involving morphologically-aware surprisal estimates and suggesting a new method for evaluating morphological prediction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models Are Human-Like Internally

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Larger LMs appear human-like internally: when surprisal is read from internal layers, their fit to human reading and EEG data matches or exceeds smaller models, overturning final-layer-only conclusions.

Pith tools