Pith. sign in

REVIEW 2 cited by

Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.14938 v2 pith:SMBNTNYC submitted 2023-05-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords benchmarklanguagellmsmodelssocialtaskssocketcategories
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have been shown to perform well at a variety of syntactic, discourse, and reasoning tasks. While LLMs are increasingly deployed in many forms including conversational agents that interact with humans, we lack a grounded benchmark to measure how well LLMs understand \textit{social} language. Here, we introduce a new theory-driven benchmark, SocKET, that contains 58 NLP tasks testing social knowledge which we group into five categories: humor & sarcasm, offensiveness, sentiment & emotion, and trustworthiness. In tests on the benchmark, we demonstrate that current models attain only moderate performance but reveal significant potential for task transfer among different types and categories of tasks, which were predicted from theory. Through zero-shot evaluations, we show that pretrained models already possess some innate but limited capabilities of social language understanding and training on one category of tasks can improve zero-shot testing on others. Our benchmark provides a systematic way to analyze model performance on an important dimension of language and points to clear room for improvement to build more socially-aware LLMs. The associated resources are released at https://github.com/minjechoi/SOCKET.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Extracting Participation in Collective Action from Social Media

    cs.SI 2025-01 conditional novelty 6.0 of 10

    A new classifier suite detects and levels expressions of collective action participation in Reddit comments, reaching weighted F1=0.71 for binary detection.

  2. QA-TOOLBOX: Conversational Question-Answering for process task guidance in manufacturing

    cs.CL 2024-12 conditional novelty 5.0 of 10

    An LLM-augmented dataset and baseline evaluation for manufacturing task guidance QA, using LLM-as-a-judge with expert validation.

Pith tools