Pith. sign in

REVIEW 1 cited by

From Millions of Tweets to Actionable Insights: Leveraging LLMs for User Profiling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.06184 v1 pith:DP3YQZZY submitted 2025-05-09 cs.SI cs.CLcs.IR

classification cs.SIcs.CLcs.IR
keywords userprofilingknowledgelargellmsmethodprofilesadaptable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Social media user profiling through content analysis is crucial for tasks like misinformation detection, engagement prediction, hate speech monitoring, and user behavior modeling. However, existing profiling techniques, including tweet summarization, attribute-based profiling, and latent representation learning, face significant limitations: they often lack transferability, produce non-interpretable features, require large labeled datasets, or rely on rigid predefined categories that limit adaptability. We introduce a novel large language model (LLM)-based approach that leverages domain-defining statements, which serve as key characteristics outlining the important pillars of a domain as foundations for profiling. Our two-stage method first employs semi-supervised filtering with a domain-specific knowledge base, then generates both abstractive (synthesized descriptions) and extractive (representative tweet selections) user profiles. By harnessing LLMs' inherent knowledge with minimal human validation, our approach is adaptable across domains while reducing the need for large labeled datasets. Our method generates interpretable natural language user profiles, condensing extensive user data into a scale that unlocks LLMs' reasoning and knowledge capabilities for downstream social network tasks. We contribute a Persian political Twitter (X) dataset and an LLM-based evaluation framework with human validation. Experimental results show our method significantly outperforms state-of-the-art LLM-based and traditional methods by 9.8%, demonstrating its effectiveness in creating flexible, adaptable, and interpretable user profiles.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PolitiSky24: U.S. Political Bluesky Dataset with User Stance Labels

    cs.CL 2025-06 conditional novelty 5.0 of 10

    PolitiSky24 provides 16,044 AI-labeled user-level stance pairs for Trump and Harris from 8,467 Bluesky users, with the labeling pipeline reporting 81% validation accuracy.

Pith tools