Pith. sign in

REVIEW 18 cited by

BERT Rediscovers the Classical NLP Pipeline

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.05950 v2 pith:GLGB4JDZ submitted 2019-05-15 cs.CL

classification cs.CL
keywords modelpipelinebertinformationadjustadvancedanalysisappear
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained text encoders have rapidly advanced the state of the art on many NLP tasks. We focus on one such model, BERT, and aim to quantify where linguistic information is captured within the network. We find that the model represents the steps of the traditional NLP pipeline in an interpretable and localizable way, and that the regions responsible for each step appear in the expected sequence: POS tagging, parsing, NER, semantic roles, then coreference. Qualitative analysis reveals that the model can and often does adjust this pipeline dynamically, revising lower-level decisions on the basis of disambiguating information from higher-level representations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 50 citations worldwide. Full citation record

  1. On the Interpretability of Whisper Encodings Using Sparse Autoencoders

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Sparse autoencoders applied to Whisper ASR reveal monosemantic features across linguistic boundaries and demonstrate cross-lingual feature steering.

  2. Epistemic Familiarity is Associated With Belief Stability in Large Language Models

    cs.CL 2025-11 conditional novelty 7.0 of 10

    Language models retract previously true answers far more often after seeing unfamiliar synthetic statements than after seeing familiar fictional statements, in both internal probes and prompted behavior.

  3. Mobius Learning: Cyclic Depth Folding in Transformers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Möbius Learning, which cyclically shifts block order across data streams, achieves lower validation loss than fixed-order looped training at loop depths 6, 10, and 15 in a 124M-parameter GPT-2 experiment.

  4. Prototype Transformer: Towards Language Model Architectures Interpretable by Design

    cs.AI 2026-02 conditional novelty 6.0 of 10

    ProtoT is an autoregressive language model whose attention is replaced by learned prototype channels that are claimed to capture nameable concepts and allow targeted edits, at linear sequence cost but with slightly lo...

  5. Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Diffusion language models can revise away harmful intermediate text, and a step-wise internal refusal signal detects jailbreaks cheaply across autoregressive and diffusion models.

  6. Geometry of Reason: Spectral Signatures of Valid Mathematical Reasoning

    cs.LG 2026-01 reject novelty 6.0 of 10

    Spectral features of attention are claimed to classify proof validity with near-perfect effect sizes, but the main evaluation relabels proofs using the classifier's own outputs.

  7. LAYA: Layer-wise Attention Aggregation for Interpretable Depth-Aware Neural Networks

    cs.LG 2025-11 conditional novelty 6.0 of 10

    LAYA replaces the last-layer classifier with an input-conditioned attention mixture of all hidden representations, yielding small accuracy gains and per-sample layer-attribution scores.

  8. AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    AIM-CoT improves multimodal chain-of-thought by selecting image regions that reduce predictive uncertainty and inserting them when attention shifts toward the visual input.

  9. All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs solve arithmetic in-context via an All-for-One pattern, with all input-specific computation occurring at the last token after a two-layer information transfer window.

  10. STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A structure-aware exemplar retriever with a hidden-state syntactic injection module improves in-context semantic parsing across four benchmarks.

  11. Meaning-infused grammar: Gradient Acceptability Shapes the Geometric Representations of Constructions in LLMs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    In Pythia language models, the geometric distance between Double Object and Prepositional Object representations grows with human-rated preference strength, evidence for graded construction representations.

  12. From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Fine-tuned LLMs extract STIX entities and relationships from threat reports with per-module F1 scores of 84.4%, 88.5%, 95.5%, and 84.6%, backed by a new 4,011-entity annotated dataset.

  13. Mechanistic Decomposition of Sentence Representations

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Sentence embeddings can be decomposed into sparse, interpretable atoms via supervised dictionary learning, and mean pooling preserves mainly atoms aligned with the sentence direction.

  14. Adaptive Task Vectors for Large Language Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Adaptive Task Vectors use a small model to generate query-specific steering vectors for frozen LLMs, reporting strong accuracy and generalization, though the theoretical equivalences to LoRA and Prefix-Tuning are not ...

  15. Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures

    cs.AI 2026-07 conditional novelty 5.0 of 10

    LLM embeddings largely reproduce the Indo-European language family tree, and how faithfully a model reproduces that tree correlates with its XNLI multilingual performance.

  16. S2Sent: Nested Selectivity Aware Sentence Representation Learning

    cs.CL 2025-08 conditional novelty 5.0 of 10

    A lightweight cross-layer fusion module, S2Sent, improves unsupervised sentence embeddings by gating and DCT frequency selection across Transformer blocks.

  17. SALMAN: Stability Analysis of Language Models Through the Maps Between Graph-based Manifolds

    cs.LG 2025-08 conditional novelty 5.0 of 10

    SALMAN ranks each text sample's fragility via the distortion between input and output embedding distances and uses the ranking to improve attack success rates and fine-tuning robustness.

  18. PrunePEFT: Iterative Hybrid Pruning for Parameter-Efficient Fine-tuning of LLMs

    cs.LG 2025-06 conditional novelty 5.0 of 10

    An iterative hybrid pruning method selects which PEFT modules to keep at each transformer layer, matching or improving fixed PEFT baselines on GLUE at 1% trainable parameters.

Pith tools