Pith. sign in

REVIEW 1 cited by

Uncovering Constraint-Based Behavior in Neural Models via Targeted Fine-Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01207 v1 pith:VYN7GYFG submitted 2021-06-02 cs.CL

classification cs.CL
keywords behaviorlinguisticmodelslanguagemodelconstraintsknowledgebiases
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A growing body of literature has focused on detailing the linguistic knowledge embedded in large, pretrained language models. Existing work has shown that non-linguistic biases in models can drive model behavior away from linguistic generalizations. We hypothesized that competing linguistic processes within a language, rather than just non-linguistic model biases, could obscure underlying linguistic knowledge. We tested this claim by exploring a single phenomenon in four languages: English, Chinese, Spanish, and Italian. While human behavior has been found to be similar across languages, we find cross-linguistic variation in model behavior. We show that competing processes in a language act as constraints on model behavior and demonstrate that targeted fine-tuning can re-weight the learned constraints, uncovering otherwise dormant linguistic knowledge in models. Our results suggest that models need to learn both the linguistic constraints in a language and their relative ranking, with mismatches in either producing non-human-like behavior.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Implicit Causality-biases in humans and LLMs as a tool for benchmarking LLM discourse capabilities

    cs.CL 2025-01 conditional novelty 7.0 of 10

    Most tested LLMs fail to reproduce human implicit causality biases in coreference, coherence, and referring-expression form, even when they show partial coreference effects.

Pith tools